I keep coming back to a specific failure I have watched repeat across enterprise agent rollouts. A team ships something ambitious. Usage climbs. The internal deck looks great. Then a renewal conversation happens twelve months later with a procurement lead who was never in the original room, and the whole thing goes quiet. Nobody can point to the artifact that would answer a simple question: what did this agent actually change that we would pay to keep?

My read is that this is not a data problem. It is not a governance problem, or a proof-loop problem, or a measurement lag problem, though all of those get named in the postmortem. It is a sequencing problem. The team started with the pipeline instead of the definition. They wired the agent into the workflow before anyone made a testable commitment about what a good outcome would look like when the human champion who signed the original contract had rotated to another role.

That sequencing error is the interesting one because it is invisible during the exciting part of the project. Agents are fluent. Dashboards fill in. Stakeholders are curious. Nothing feels wrong until a stranger in a renewal room has to reconstruct the business case from artifacts the system was never designed to emit.

The order matters more than the stack

When I have seen agent programs land well, the first design decision was not the retrieval architecture, the model choice, the eval suite, or the integration path. It was a written answer to two questions the team was willing to be held to. What does this agent make measurably better for the business, stated in terms a non-user could verify? And what specific artifact will the workflow produce, on its own, that would let someone reconstruct that answer without interviewing anyone?

Those two questions are the objective function and the acceptance rubric. They sound almost boring. They are also the thing most teams skip, because skipping them lets you ship faster and because the reward for asking them shows up much later, in a room you are not in.

I would rather have a mediocre pipeline with a sharp definition of good than a beautiful pipeline with a fuzzy one. The mediocre pipeline can be improved against a fixed target. The beautiful pipeline can only be defended by the person who built it, and that person is not in the renewal meeting.

What the frontier firm framing actually implies

Microsoft's 2026 Work Trend Index writeup frames the shift as four collaboration patterns, from human as producer to human as orchestrator of parallel agents. Most of the coverage treats this as a maturity ladder. That is not the useful reading.

The useful reading is that as you climb the ladder, the human job changes from doing the work to defining what the work should produce and evaluating whether it did. Microsoft is direct about this. As agent activity rises, the tactical execution shrinks and the need to set direction, define standards, and evaluate outcomes rises. Their own survey data has AI users naming quality control of AI output and critical thinking as the two most important human skills in that new configuration.

Read that literally. If leadership is now the standard-setting and evaluation layer, then the standard and the evaluation criteria are the leadership deliverable. They are not something you can retrofit after the agent is running. They are the first artifact of the program, and everything else, including the data plumbing, is downstream engineering against them.

This is why the sequencing matters. An orchestrator without a definition of good is not orchestrating anything. It is just running fast.

The renewal-grade artifact

I have written before about hollow usage, where agents become the primary user of a product and the human champion who used to translate activity into business value quietly disappears from the account. That post named the problem. This one is about the leadership move that prevents it.

The move is to design, before deployment, the specific artifact the workflow will emit that would convince someone who was not in the original buying conversation. Not a dashboard. Not a usage chart. An artifact that shows an outcome the buyer actually cares about and ties it, credibly, to something the agent did.

For a sales agent that reduces admin, the artifact is not seat activity. It is reclaimed seller hours mapped to specific customer-facing behaviors that changed, and the pipeline or relationship outcomes those behaviors produced. For a support agent, it is deflection tied to satisfaction and cost per resolved case, with the harder cases visibly escalated rather than quietly closed. For a marketing agent, it is claims shipped with provenance and the campaign outcomes those claims moved.

The test is simple. If the agent stopped running tomorrow, could you show a stranger, in one page, what would be worse? If the answer requires the person who built the program to narrate it, you do not have an artifact. You have a story that lives in one head.

Designing that artifact up front changes what you build. It changes what you log, what you surface, what you refuse to promise, and what you insist the workflow produce as a first-class output rather than an analytics afterthought. This is the same discipline I described in when the loop runs itself and acceptance becomes the job, pushed one step earlier: acceptance criteria are not a launch gate, they are a design input.

The uncomfortable part

The reason teams skip this is not laziness. It is that writing a testable definition of good is a commitment. You are saying, in advance, what you will and will not claim credit for. You are giving up some of the interpretive flexibility that makes internal storytelling easier. You are constraining the program to outcomes the workflow can actually produce evidence for, which means saying no to ambitions the workflow cannot support yet.

That restraint is the leadership act. It is unglamorous compared to picking a model or announcing a partnership. It also compounds. A program with a clear objective function and a defensible artifact renews at higher confidence, expands into adjacent use cases on the strength of a track record rather than a pitch, and gives the next buyer, who did not sign the first contract, something to point at.

A program without those things renews on relationships, which is fine until the relationship rotates.

What this makes possible

The interesting thing about getting the sequence right is that it turns agent speed into pricing power instead of discount pressure. When the workflow emits its own evidence, you are no longer negotiating from usage. You are negotiating from a documented change in the customer's business, produced by the product, on a cadence the customer can verify. That is a very different conversation.

It also changes how the internal team spends its time. Engineering effort moves from post-hoc measurement scaffolding to designing the workflow so the evidence falls out of normal operation. Product moves from feature lists to outcome commitments. Go-to-market moves from proof-by-anecdote to proof-by-artifact. None of that requires a bigger model. It requires deciding what you are trying to prove before you start proving it.

The frontier firm posture Microsoft describes is only useful if leaders take the direction-setting and evaluation part seriously as their actual job. Orchestration is not a tooling upgrade. It is a demand that leaders make the definition of good precise enough to hand to a system.

The question to carry into the next approval

The next time an agent rollout comes across your desk for sign-off, before the demo, before the architecture review, before the rollout plan, ask one question and refuse to move on until it has a specific answer.

What artifact will this workflow produce, without anyone narrating it, that would convince your counterpart in a renewal room twelve months from now that this was worth keeping?

If the team can answer that in a sentence, the rest of the program has a spine. If they cannot, you are not approving an agent. You are approving activity, and activity is not what gets renewed.

Sources