Skip to content
Impulse TeamsImpulse Teams

News

The AI Baseline Moved

July 28, 2026

Grainy gold light pressing across a dark abstract field beneath a centered AI BASELINE label

In January 2025, OpenAI introduced computer use as a research preview. Its Computer-Using Agent could see a screen and operate a mouse and keyboard, but OpenAI reported a 38.1% success rate on OSWorld against a 72.4% human reference. The capability was real, useful in narrow cases, and visibly early.

Eighteen months later, the unit of AI work looks different. OpenAI says ChatGPT Work can act across apps and files for hours. Anthropic says Claude Code can coordinate hundreds of parallel subagents and verify their output. xAI positions Grok 4.5 for coding, agentic tasks, and work inside Excel, PowerPoint, and Word. These are vendor claims, not independent proof of reliable business outcomes. They still mark a concrete shift: frontier AI is moving from producing answers toward carrying work through tools.

The change is larger than a better model

Model capability is only the first layer.

The platform around the model has expanded at the same time. In April 2026, Google launched Gemini Enterprise Agent Platform with long-running state, persistent memory, agent identity, registry, gateway controls, simulation, evaluation, and observability. These are the components needed to run agents as managed systems instead of isolated demonstrations.

Products are also exposing more of that stack. Claude Code became generally available in May 2025. One year later, Anthropic's dynamic workflows research preview could split large work across parallel agents. In July 2026, OpenAI presented ChatGPT Work as a cross-application agent that can produce documents, sheets, slides, and web apps. xAI presented Grok 4.5 as both an engineering model and an office-work tool.

None of this proves that every agent can be trusted with every workflow. It shows that the available design space changed. A system built around one user, one prompt, and one response may still function while exposing much less capability than current platforms and products can support.

Five layers are moving at different speeds

The new baseline does not reach a business in one step. It moves through five layers:

  1. Model capabilities: reasoning, multimodal work, tool use, speed, cost, and reliability determine what is technically possible.
  2. Platform capabilities: runtimes, memory, permissions, evaluations, and observability determine what can be deployed and controlled.
  3. Product, SaaS, and tooling capabilities: the software a team buys determines which underlying capabilities are actually exposed.
  4. Business-system adoption: workflows, data access, integrations, ownership, and review rules determine whether the capability changes real operations.
  5. Employee adoption and enablement: people need to know what to delegate, how to inspect the result, and when to intervene.

The first three layers can change in a release cycle. The last two require decisions, system work, changed habits, and evidence from live use. That difference in speed is the capability gap.

Availability is not evidence of business value

Microsoft's 2026 Work Trend Index gives a useful, bounded view of the lower layers. Microsoft reports 15 times year-over-year growth in active agents across its Microsoft 365 agent ecosystem from March 2025 to March 2026. It does not publish the absolute counts.

The same report surveyed 20,000 knowledge workers who already use AI across ten markets. Microsoft classified 19% as having both high individual and organizational readiness, while 10% reported strong individual readiness but weak organizational support.

Those results should not be treated as causal proof. The readiness measures and reported outcomes are self-reported, non-users were screened out, and Microsoft explicitly describes the relationships as statistical associations rather than causal effects. The evidence supports a narrower conclusion: access and activity can rise quickly while operating conditions remain uneven.

That is why a vendor release is not an adoption result. A company captures the gain only when the new capability survives its data, permissions, quality standards, handoffs, incentives, and day-to-day use.

Build an upgrade path, not a release habit

Businesses do not need to adopt every new model or feature. They need a repeatable way to notice when the baseline has moved enough to justify change.

  • Track capability transitions, not release volume. Ask what can now be completed, controlled, or recovered that could not be handled reliably before.
  • Test the current system against real work. Compare quality, cost, completion time, review effort, and failure recovery with the existing baseline.
  • Locate the slowest layer. A stronger model will not fix missing permissions, poor data, an unsupported product, an ownerless workflow, or employees who do not know how to review the result.
  • Upgrade the affected system, not only the model. Change the workflow, controls, integrations, responsibilities, and enablement needed to make the capability usable.
  • Leave the system unchanged when the improvement is not material. Selective upgrades are a discipline. Constant switching is not.

The strategic question is no longer whether a company has access to AI. It is how much time passes between a meaningful capability change and a reliable business result.

That elapsed time is The Capability Gap.

Related services: AI strategy, Agent implementation, Enablement

Sources

Ready to build your own update?

Tell us your current blockers and desired outcomes. We will propose a practical first execution scope.