I keep coming back to one number from IBM: 11% readiness.

That is the share of CIOs and CTOs who told IBM's Institute for Business Value they feel fully prepared for the scale of AI agent deployment expected in the next year. The same study puts a sharper edge on it: two-thirds of tech leaders say they are accountable for AI systems they do not fully control, and 70% report that teams across the business are deploying faster than IT can track. IBM calls it a control gap. By 2027, they expect agent deployments to grow 38%.

Most coverage will read this as a people problem. Leaders are behind. Boards need more literacy. Change management is lagging.

My read is different. The 11% is not a readiness verdict on humans. It is a readiness verdict on the products those humans are being asked to buy.

If you are accountable for something you cannot observe, the missing piece is not your courage. The missing piece is the instrument panel. Agent vendors have been shipping autonomy without shipping the control surface that makes autonomy legible. That is a product deficiency, and it is the most interesting opening in the enterprise stack right now.

Three primitives that should already be in the box

When a technology leader says they are not ready, ask what artifact would make them ready. In almost every conversation I have had this year, the answer collapses into three product primitives that agent vendors should be shipping as first-class features, not stitching in after the pilot.

Decision provenance. For any action an agent takes, the system should emit a compact, queryable record of why. Which tools were called. Which policy was consulted. Which data version was in scope. Which prior state triggered escalation. Not a log dump. A structured trail that an auditor, a lawyer, or a line manager can read without a data science team in the loop.

Audit artifacts. Provenance is the raw material. Artifacts are the finished object. A weekly rollup that a CFO can sign. A per-workflow evidence package that maps agent actions to controls the company already runs. The vendor's job is to produce this, not to hand the customer a firehose and call it transparency.

Refresh cadence. Agents drift. Models update. Connectors change. Policy shifts. A serious product tells you when its assumptions have gone stale, forces a re-approval, and versions the whole boundary set the same way software teams version releases. Static governance documents cannot chase live systems. The product has to carry the clock.

None of this is exotic. It is the enterprise version of what already exists in observability, financial audit, and safety-critical software. What is new is treating it as a shipping requirement for autonomy, not a services engagement after go-live.

Why the gap sits at the product layer, not the person layer

The more interesting inference from the IBM numbers is that the 11% is durable. You cannot train your way out of it. The people are not going to catch up to a system that emits no evidence.

This is where the diagnosis matters. If you accept the readiness framing, you buy more workshops. If you accept the product framing, you rewrite the RFP. I have watched enough agent evaluations this year to have a strong opinion about which one moves the needle.

The control gap is what happens when autonomy scales faster than the instruments meant to observe it. IBM's own vendor articulation of an AI operating model treats policy enforcement and multi-agent auditability as infrastructure. Snowflake has been arguing something similar with its control-plane framing for the agentic enterprise. The vendors closest to the problem are converging on the same insight from different sides: coordination and evidence have to live inside the product.

This is also where the productivity measurement problem and the integration debt problem meet. If the agent cannot show its work in a form the business can accept, the workflow stalls at the last human review. The bottleneck migrates from model quality to the seam between the model and the decision it feeds.

The Australia analogy, one paragraph only

A useful cautionary parallel sits outside the AI beat. Australia doubled penalties on platforms that fail to enforce its under-16 social media ban after evidence that teens were routing around it with fake mustaches for age checks and VPNs. A study of 400 children found little change in usage. Researchers concluded the policy would probably not deliver the mental-health benefits its supporters expected. The lesson translates directly: enforcement that is not embedded in the product becomes theater. Bolt-on AI governance will end the same way. If control is not native to the agent, it will not survive contact with a workflow that is measured on speed.

What I ask agent vendors now

I keep the same short list in every vendor conversation, and it has done more work than any capability demo I have sat through.

Show me the decision record for a single agent action, end to end, without a screenshare of a database. Show me the artifact you would hand a nervous auditor. Show me what happens when the underlying model version changes on Tuesday. Show me the stop condition, not as a promise, but as a configurable object I can inspect. Show me the refresh cadence for policy, and how the system tells my team it is time to re-approve.

If a vendor cannot answer those in the demo, they are not selling an enterprise agent. They are selling a script that runs unattended. Those are different products, and they belong in different budget lines.

The stop contract work I have written about before sits inside this same shift. A legible boundary between agent action and human escalation is not a feature you add. It is a product primitive that decides whether the workflow can be trusted to run at all.

The buying implication

If you take the IBM number seriously, it changes how you spend in 2026.

Weight vendor selection toward products that emit control as they act. Score decision provenance, audit artifacts, refresh cadence, and configurable stop conditions with the same seriousness you give accuracy benchmarks and latency. Ask for the evidence object, not the trust claim. Reject anything that requires your team to build the observability layer after signature. Budget refresh, not just deployment. Assume any static policy document is stale within a quarter, and buy products that treat that as their problem to solve.

The 11% is not a call for humility. It is a specification you have not yet written into your RFP.

Agents will move at AI speed regardless of whether the enterprise is ready. The only question worth arguing about is whether the products they run on emit enough evidence for a human being to assent to what just happened. That is what closes the gap. Not more education. Not more governance decks. A shorter distance between the action and the artifact that proves it was the right one.

Buy for that, and 11% becomes an artifact of last year's product generation.

Sources