Domain Expertise Is the Agent Multiplier. Institutional Memory Is the Interface.
Across roughly 400,000 Claude Code sessions, Anthropic found domain expertise predicts how much useful work an agent produces per instruction. OpenAI’s Codex adoption shows the same bottleneck moving into departments with real domain fluency. When execution gets cheaper, power shifts to what your o…

Anthropic published a quiet but important finding this fall: across roughly 400,000 Claude Code sessions, the variable that best predicted how much useful work an agent produced per instruction was not coding skill. It was domain expertise.
That is a smaller signal than the headline would suggest, and easy to overclaim. One company, one tool, one seven-month window. But it lands next to OpenAI's own data on Codex adoption, and together they describe a real shift in where advantage sits inside AI-native work.
The temptation is to read this as another dispatch about agents eating knowledge work. My read is narrower and more useful. Execution is being absorbed. What isn't being absorbed, and what compounds when everything else gets cheaper, is the human capacity to specify what matters and the organizational capacity to remember what is true.
What the data actually says
Anthropic's analysis of Claude Code usage frames the pattern plainly: in typical sessions, humans make most of the planning decisions and the agent makes most of the execution decisions. When users bring more domain knowledge, they plan more, the agent executes more, and sessions are more likely to finish successfully end to end. The multiplier is not tool fluency. It is the quality of direction.
OpenAI's companion account of Codex adoption sharpens the picture. Inside OpenAI, Codex accounts for 99.8% of weekly output tokens. Legal, Finance, and Recruiting became Codex-primary around April 2026. By May, more than 80% of sampled individual users had made at least one request estimated to exceed thirty minutes of human work, and a quarter had made one estimated to exceed eight hours.
Two things are happening at once. Agents are absorbing longer horizons of execution. And they are moving into departments where the users have deep domain fluency and no coding background. Recruiting professionals directing agents through a coding tool. Lawyers directing agents through a coding tool. The people whose planning quality is highest for their domain get the largest lift.
That is the pattern worth taking seriously. It is not a claim about model quality. It is a claim about where the bottleneck moved.
Execution is not the scarce resource anymore
When execution becomes something you delegate in a paragraph of instruction, the constraint slides upstream. It stops being "can we do this?" and becomes "can we say precisely what we want done, and can we recall the context that makes the answer correct?"
Say precisely and recall accurately. Those two capacities are not model features. They are organizational ones.
The first is domain fluency: the tacit knowledge of what a good answer looks like in your specific problem space, why a plausible output is actually wrong, which edge cases matter. The Anthropic data suggests this is a real, measurable multiplier rather than a soft skill. Users with more domain depth extract more output per instruction because they can specify constraints an outsider would not know to write.
The second is institutional memory. That is where the practical opportunity opens up.
Memory as an agent-facing interface
Most organizations already have institutional memory. It lives in slack threads, in the heads of six people, in decks nobody can find, in decisions half-remembered from two reorgs ago. It is not retrievable at machine speed. It is not consistent enough to steer an agent against.
The move I keep coming back to is treating institutional memory as a product surface rather than a folder. Origin narratives, decision logs, canonical facts, prior tradeoffs, the specific reasons a project failed in 2022. All of it curated to be retrievable by an agent with a clear provenance trail.
A low-drama example. Liverpool Football Club is 134 years old. Every match, transfer, tactical debate, and ownership decision sits inside a continuously curated record. When commentators reach for context, the record answers immediately. That is what an institutional memory interface looks like from the outside: consistent, canonical, retrievable, alive.
Most enterprises do not have that. They have archives. The difference matters because agents can only steer against memory they can query. An agent asked to draft a partnership proposal cannot infer why the last three partnerships were structured the way they were. It can guess. Its guess will be plausible and often wrong.
Microsoft's framing around work intelligence gestures at this. So does the observation sitting inside OpenAI's Frontier positioning that agents get deployed everywhere and each one is isolated in what it can see and do. The gap between model capability and delivered outcome is context, and context is memory made retrievable.
Where I would put the next dollar
The dollar most enterprises are spending right now is on tool access. Seats, subscriptions, pilots. That spend produces motion. It does not produce the multiplier the Anthropic data describes.
My read is that the next dollar belongs in two places, and neither is a tool.
The first is domain fluency, treated as a first-class investment. Not generic AI training. Training on the actual problem space, on what good looks like in your industry, on the constraints an agent will not infer. This is what makes a user in Legal or Recruiting suddenly get large output gains from a coding agent. It is why the domain expert multiplier shows up in the data at all.
The second is institutional memory infrastructure. Canonical records. Decision logs written to be queried. Provenance metadata attached to statements of fact. A retrieval interface an agent can plan against and a human can audit. This is closer to what I've called integration debt, pointed further upstream. Beyond fixing the seam where AI output meets decisions, the work is building the memory layer that lets agents propose good decisions in the first place.
Neither investment shows up on a tool procurement line. Both compound.
The shift in what leaders own
There is a version of the acceptance argument I've made before: when the loop plans and executes, the human role becomes defining what counts as ready. The Anthropic data extends that argument by naming the specific capacity that makes acceptance possible. It is domain fluency, held either in the person directing the agent or in the memory the agent can query.
This is a quieter power shift than the "AI transforms everything" pitch, and it distributes advantage in a specific way. Organizations that own deep, curated, retrievable knowledge of their domain will get compounding leverage from every capability release. Organizations that treat AI as a productivity layer on top of an unchanged memory stack will keep buying velocity and wondering where the returns went.
The thing I would watch is whether leadership teams start allocating to fluency and canonical records with the same seriousness they allocate to tool access. My guess is that the ones who do will look, two years out, like they got lucky with a model. They didn't. They invested in the input that actually predicted output.
The question stops being whether your team can use agents. It becomes whether your organization can reliably remember what is true, and whether it can hand that memory to an agent in a form the agent can act on. Everything else is motion.
Sources and further reading
- How agents are transforming work. A concrete, vendor-backed pattern showing agents are changing the default unit of organizational work, beyond improving model quality. That implies a near-term business opportunity: redesign knowledge workflows around delegated runs that plan, act, iterate, and finish, rather than around single-shot responses.
- OpenAI Frontier. Historically indexed source that semantically overlaps with the current article frame.
- Introducing the First Frontier Suite built on Intelligence + Trust. Historically indexed source that semantically overlaps with the current article frame.
- Integration Debt: Why AI Output Fails at the Decision Layer. Published AI Stoic archive memory that supports crosslinking and non-repetition.
- When the Loop Runs Itself, Acceptance Becomes the Job. Adjacent published post that supports internal crosslinking.
Reader account
Join the conversation
Sign in with a private email link to manage preferences and leave a comment.

Comments