I keep coming back to a mismatch I can feel in every institution that says it wants AI and then quietly refuses to let it decide anything.

The AI layer is probabilistic. It generates. It samples. It varies with context, prompt, and version. That is not a flaw. It is the reason the technology is interesting. But the layer sitting on top of it, the one procurement, audit, general counsel, and boards actually operate, is deterministic by design. It expects a stable artifact. Something a person can sign. Something a regulator can subpoena. Something a judge can point to and say, this is what was decided, and this is who decided it.

Most organizations pretend this gap does not exist. They write a policy. They stand up a committee. They add another slide to the AI governance deck. And then they wonder why their pilots keep stalling in the last mile between demo and deployment.

My read is that the fix is not another policy. It is a product.

A judicial platform that names the boundary

The cleanest current example is not a McKinsey deck. It is a courtroom.

Semafor Gulf reported that the Abu Dhabi Judicial Department has deployed a fully integrated AI platform that assists judges with case analysis, legal research, and drafting. Judges keep the final ruling. The platform runs at operational scale, not as a demo. The architecture is explicit: assist is probabilistic, decide is deterministic, and the interface between them is engineered rather than assumed.

That is not a minor design detail. It is the whole thing.

A judicial system is the hardest possible test case for AI adoption because it cannot tolerate the failure modes we tolerate in marketing copy or customer support. A ruling has to be stable. It has to be attributable. It has to be defensible on appeal against an adversary who is paid to find every crack in the reasoning. If AI is going to sit inside that workflow at all, someone has to have solved the problem of what the machine hands to the human, and what the human hands to the record.

That handoff is the product.

What an evidence contract actually is

I have been calling this pattern an evidence contract, and I wrote about it earlier through the lens of the EU's Entry/Exit System. The border case showed the failure mode: probabilistic execution slamming into deterministic institutional gates and generating queues. The judicial case shows the working version. Same physics, different design maturity.

An evidence contract is not a policy document. It is a set of structured artifacts that a probabilistic system produces on every run, engineered so that a deterministic institution can accept them without translation.

Concretely, that means at minimum:

  • The inputs the system saw, captured verbatim, versioned, and timestamped. Not summarized. The actual prompt, retrieved documents, and tool state.
  • A decision trace that shows which candidate outputs were considered, which were rejected, and on what basis. This is not chain of thought theater. It is the record a reviewer needs to understand why the output looks the way it does.
  • Citations and provenance for any factual claim, tied to source identifiers a receiving system can independently verify.
  • Guardrail records showing which policies fired, which were passed, and where a human was called in.
  • A named human acceptance point, with role, timestamp, and the version of the artifact they accepted.

Get that stack right and the probabilistic system stops looking like a black box to the institution around it. The model can still be doing whatever it is doing internally. What leaves the system is a package a lawyer, an auditor, or a judge can hold in their hand and reason about the same way they reason about any other institutional artifact.

What leaders can control is the shape of that package. They cannot control the model's internal path. That distinction is the entire game.

Why the checklists keep failing

Most AI governance work I see in the wild is not solving this. It is solving something adjacent and easier: producing documents that make committees feel governed.

Risk registers. Model cards. Vendor questionnaires. Ethical AI principles pinned to a wall. All of it useful at the margin, none of it addressing the actual bottleneck, which is that when the AI hands its output to the next system or the next person, that receiving surface does not know what to do with a probabilistic artifact. It was built to accept a signed one.

You cannot policy your way across that gap. The receiving system does not read policies. It reads fields.

This is why the interesting AI adoption stories are converging on the same insight: governance is an operating-model component, not a slow approval layer. It is something you build into the workflow at the artifact level, not something you bolt on at the review level. The teams doing this well are treating the audit-ready output as a feature of the product, shipped by the same engineers who ship the model call.

There is a parallel legal debate about handling this on the accountability side rather than the design side. Some scholars have proposed regulating frontier AI like blasting or wild animals, with strict liability and mandatory insurance instead of negligence standards. That is a real conversation, and it may end up mattering. But strict liability is a settlement mechanism after harm. An evidence contract is a design mechanism before it. Institutions that ship the second do not need to litigate as much of the first.

Governance as a product primitive

Here is what I would watch for over the next year, especially among leaders who actually intend to move AI into decision workflows rather than talk about it.

The organizations that will pull ahead are the ones treating evidence design as a core product capability. Not as a legal deliverable. Not as a compliance function. As a shipped artifact, owned by product and engineering, refined by legal and audit, and specified as tightly as any customer-facing feature.

That sounds mundane until you compare adoption rates. When OpenAI studied how enterprises actually scale AI, the pattern that kept showing up was that governance-as-enabler, brought in early as a design partner, correlated with faster deployment and fewer reversals. Not slower. Faster. Because the receiving institution knew what it was accepting.

The organizations still stuck in permission loops are the ones treating governance as a review gate on top of a finished system. They are asking committees to approve outputs the committees have no structured way to evaluate. Of course they say no. The refusal is not conservatism. It is the correct response to being handed something that does not fit any existing acceptance schema.

The fix is upstream. Design the schema. Ship the artifact. Make the probabilistic step legible to the deterministic system that has to sign for it.

This is also, incidentally, why the human signature stays fixed in every serious AI workflow I have looked at. Courts have been clear. Regulators are converging. The point of the evidence contract is not to replace that signature. It is to give the human signing it something they can actually reason about, so the signature means what it used to mean.

The closing question

Abu Dhabi did not solve institutional AI adoption by having better models than everyone else. It solved a design problem everyone else is still treating as a debate.

The leaders I trust on this are not the ones with the most sophisticated views on AI safety. They are the ones who can point to the specific artifact their institution accepts, name the fields it contains, name the human whose role is to sign it, and describe what happens if it fails an audit two years from now. That is a shippable answer. Everything else is a working group.

If you cannot name the evidence artifact your institution can stand behind, you do not yet have AI governance. You have AI discussion.

Governance-as-product is the competitive advantage hiding inside the governance conversation, and most organizations will miss it because they are still treating it as a cost.

Sources