Boards still ask for AI demos. CFOs and defense planners are already funding systems that can close books, compress licensing cycles, and survive audit. The distinction matters more than most strategy decks admit.

The visible AI story — the one that dominates earnings calls and conference keynotes — is still interface-led. A chatbot answers questions. A copilot drafts an email. An image generator produces marketing assets. These are real products with real users, but they are not where the largest, most durable institutional budgets are forming.

The budgeted story is different. It is workflow-led. It lives in approval queues, reconciliation engines, document-conversion pipelines, and program-of-record designations. It is not optimized for impressions or engagement. It is optimized for cycle-time compression, backlog clearance, audit coverage, and cost-of-delay reduction inside institutions where friction is structural, not incidental.

The thesis of this piece is direct: AI's first durable breakout market is institutional administrative throughput — agents embedded in governed workflows across finance, regulation, and defense that are budgeted, auditable, and mission-critical. Three named deployments, operating across three different sectors, are already proving this thesis. And the economic value they are unlocking — not the software revenue they generate, but the friction they remove — is conservatively bounded in the hundreds of billions, with a credible path to a trillion-dollar value pool within this decade.

Follow the approval queue, not the interface.


Procurement proof before market size

Before making a large market claim, I want to set the evidentiary standard clearly. Too many AI market-sizing exercises start with a revenue forecast and work backward toward justification. That is the wrong sequence for boards that need to make capital allocation decisions.

A credible signal of durable AI deployment meets four conditions:

  1. Named. The deployment is publicly attributable to a specific institution, with identifiable decision-makers.
  2. Funded. It sits inside a core operating budget or a formal program of record — not an innovation sandbox or a pilot fund.
  3. Embedded. It operates within a core workflow, not adjacent to one. It touches transactions, approvals, or submissions that the institution cannot function without.
  4. Measurable. It is tied to KPIs the institution already tracks: days-to-close, exceptions per analyst, submission-cycle time, decision latency, or cost-of-delay.

Programs of record, core operating budgets, and formal oversight transfers matter more than pilot announcements or demo narratives. If a deployment has a budget line, an owner, and an audit trail, it is real. If it has a press release and a proof-of-concept, it is aspirational.

One more distinction matters for intellectual honesty: software spend is not the same as economic value removed. When an AI agent compresses a four-week licensing process into a single day, the software license fee is a small fraction of the economic value created. The carrying cost of a stalled nuclear project, the working-capital impact of a faster financial close, the reduction in compliance exposure from better audit coverage — these are the governing economic units. This essay measures the market in terms of friction removed, not vendor revenue captured. Confusing the two is how trillion-dollar claims become unserious.


Dell: reconciliation as the first workflow archetype

The first archetype is the one closest to most boards' own operating reality: enterprise finance.

Dell's CFO David Kennedy has built something that looks less like an AI experiment and more like an organizational redesign. As reported by Fortune, Kennedy embedded a team of data scientists directly inside Dell's finance function — not in a separate innovation lab, not in IT, not in a center of excellence that reports to a chief strategy officer. Inside finance. Reporting to finance leadership. Measured on finance KPIs.

The agents this team has built handle reconciliations, journal entries, and forecasting — the core transactional infrastructure of a finance function that supports a company generating $25 billion in AI infrastructure revenue in two years. These are not experimental use cases. They are load-bearing.

What makes this case structurally significant is the organizational design choice behind it. The dominant enterprise AI pattern is still center-out: a central team builds tools, then tries to push adoption into business units that did not ask for them. Kennedy inverted this. Finance owns its own agents. Finance owns its own data scientists. The embedding is the strategy.

But embedding agents in a finance function does not mean removing humans from it. The human approvals that remain are precisely the ones that boards, auditors, and regulators require:

  • Exception handling. When a reconciliation fails — when the numbers do not match across systems — a human analyst investigates, categorizes, and resolves.
  • Policy sign-off. Journal entries that exceed defined thresholds, or that touch sensitive accounts, require human approval before posting.
  • Close ownership. The financial close is still owned by a named human. The agent accelerates the work; it does not certify the result.
  • Audit review. External and internal auditors review both the agent output and the governance framework around it.

The KPIs boards understand here are well-established: days-to-close, exceptions processed per analyst, forecast-cycle time, and working-capital impact. An agent that reduces days-to-close from five to three does not just save labor. It accelerates cash-flow recognition, improves working-capital positioning, and compresses the window of exposure during which financial data is in-process and therefore vulnerable to error or misstatement.

Dell's $25 billion in AI infrastructure revenue is relevant context — it signals the scale of operations this finance function must support — but it is not evidence of the software TAM for administrative throughput. The evidence is in the operating model: a finance function that redesigned itself around agents, kept humans in the approval chain, and measures results in the language of the CFO, not the CTO.


DOE: document conversion as the second workflow archetype

The second archetype moves from finance to regulation, and from transactional throughput to document-conversion throughput. The underlying economics are different, but the structural pattern is the same.

The U.S. Department of Energy and Idaho National Laboratory recently used Everstar's Gordian AI to convert a Preliminary Documented Safety Analysis — a document required under DOE's authorization pathway for advanced reactor demonstrations — into sections equivalent to a U.S. Nuclear Regulatory Commission license application for commercial deployment. The tool performed semantic ontology mapping across regulatory frameworks, converting content structured for one regulatory regime into the format, logic, and compliance language of another.

The named metric is striking: a 208-page document generated in one day, versus a four-to-six-week manual process involving a team of specialists. That is not a marginal productivity improvement. It is a category change in throughput.

But the metric alone is not what makes this case credible. What matters is what happened next. The DOE confirmed that Gordian's output was subsequently evaluated by a subject-matter expert for accuracy, missing information, consistency, grammar, and structure. The expert review found that the output demonstrated quality, rigor, and depth — and, critically, that the tool could identify and qualify its own gaps in data knowledge. The human was not removed from the workflow. The human was repositioned from production to review — from doing the conversion work to certifying it.

The economics of this archetype are governed by cost of delay. A nuclear technology company pursuing commercial deployment is carrying the full cost of its project — capital, engineering, personnel, regulatory counsel — during every week that its licensing submission is not complete. If the licensing document-conversion step alone takes four to six weeks and can be compressed to one day, the value created is not the cost of the labor saved. It is the carrying cost of the stalled project during those lost weeks, plus the earlier revenue recognition from faster time-to-commercialization.

Deputy Assistant Secretary for Nuclear Reactors Rian Bahran framed the significance directly: this represents "more than incremental 'uplift' improvements" — it has the potential "to transform how industry prepares its regulatory submissions and deploys nuclear energy while upholding the highest standards of safety and compliance."

The DOE also noted that the current nuclear licensing process involves multiple rounds of manual document reviews and minor clerical adjustments that can take years to complete. A tool that compresses the document-production phase is not automating the regulatory decision. It is clearing the throughput bottleneck that precedes the decision, so that regulators can review complete, well-structured submissions faster.

This is the document-conversion archetype: rules-bound, document-heavy workflows where the bottleneck is not judgment but production — and where clearing production backlog has outsized economic consequences.


Maven: program-of-record status as the third workflow archetype

The third archetype is the one that signals what happens when AI stops being experimental and becomes institutional infrastructure.

As reported by Reuters, Deputy Secretary of Defense Steve Feinberg signed a policy memo designating Palantir's Maven Smart System as an official program of record. This is not a procurement decision in the conventional sense. It is a formalization decision. It enters Maven into the Future Years Defense Program as a protected budget line item, giving it visibility and stability across budget cycles that experimental programs lack.

The operational details are worth reading carefully:

  • The U.S. Army will manage all Maven contracts going forward.
  • Oversight transfers from the National Geospatial-Intelligence Agency to the Chief Digital and AI Officer within 30 days.
  • Program-of-record status is expected before the close of fiscal year 2026.

Each of these details carries institutional weight. Contract consolidation under a single service branch means procurement accountability. The oversight shift to the Chief Digital and AI Office means governance is formalized — there is a named owner with authority, budget, and reporting obligations. And the fiscal-year deadline means this is not an aspiration; it is a scheduled institutional commitment.

Maven started in 2017 as an effort to apply machine learning to drone surveillance footage. Google was the original technology partner and withdrew in 2018. What is happening now is categorically different from that origin: a system that was once experimental tooling is becoming core infrastructure with multi-year funding and cross-branch integration.

The right KPI lens here is not the one typically applied to defense AI in media coverage. It is not about autonomous weapons or science-fiction scenarios. It is about analyst throughput, decision latency, workflow consolidation, and funding stability. The program-of-record designation is evidence that the Department of Defense evaluated Maven on institutional criteria — budget defensibility, governance adequacy, operational continuity — and concluded it met the bar for formalization.

Reported investment has grown from $480 million in 2024 to $13 billion, a trajectory that only makes sense if the system is delivering measurable operational value across the combatant commands where it is deployed. That kind of budget escalation does not survive the Pentagon's internal review process on the strength of demos.


Three different sectors, one operating-model pattern

Dell's finance function, the DOE's licensing pipeline, and the Department of Defense's intelligence workflow operate in different sectors, under different regulatory regimes, with different risk profiles. But they share a common operating-model pattern that boards can use as a diagnostic lens for their own institutions.

Archetype 1: Reconciliation and exception clearing. Core finance transactions — reconciliations, journal entries, forecasting — are processed by agents, with humans retained for exception handling, policy sign-off, close ownership, and audit review. The governing metric is cycle-time compression in the financial close.

Archetype 2: Document conversion and gap detection. Rules-bound, document-heavy regulatory workflows are accelerated by converting content across regulatory frameworks, with humans repositioned from production to expert review. The governing metric is submission-readiness time and cost of delay.

Archetype 3: Routing, triage, and workflow consolidation. Multi-source intelligence and operational data are consolidated into a single architecture, with humans retained for decision authority and oversight. The governing metric is analyst throughput and decision latency across the enterprise.

All three share four structural characteristics:

  1. Document density. The workflows are document-heavy, with value locked in structured and semi-structured content.
  2. Approval gates. Human approvals are required at defined control points — not as a courtesy, but as a regulatory or operational necessity.
  3. Systems-of-record dependence. The workflows operate on or integrate with systems of record that are the institution's authoritative data sources.
  4. High cost of delay. Every day a workflow remains in-process carries a measurable economic cost — in carrying charges, compliance exposure, working-capital drag, or operational risk.

The real product category emerging here is not an interface. It is not a chatbot. It is not a copilot. It is governed throughput embedded in institutional process — the ability to clear backlogs, compress decision cycles, and improve audit coverage inside workflows that are budgeted, auditable, and mission-critical.

Executives should stop asking where to demo AI and start asking where approval cycles, audit trails, and document conversion are throttling growth. That is where defensible value is appearing first.


The control layer is not overhead — it is the product

Every workflow archetype described above shares a requirement that most AI market analyses treat as a footnote: governance. In practice, the control layer is not overhead. It is the product. Without it, agents cannot touch money, regulatory submissions, or classified data. With it, they become institutional infrastructure.

The control primitives required for agents operating in governed workflows include:

  • Role-based access control. The agent has permissioned access to specific systems, accounts, and data — not general access to everything.
  • Segregation of duties. The agent cannot both initiate and approve its own actions. Approval authority remains with a human or a separate system.
  • Immutable audit logs. Every action the agent takes is logged in a tamper-evident record that auditors and regulators can review.
  • Provenance tracking. The source of every input, the model version used, and the decision path taken are recorded and retrievable.
  • Exception routing. When the agent encounters an input outside its defined parameters, it escalates to a human — it does not guess.
  • Model change control. Updates to the underlying model are versioned, tested, and approved before deployment into production workflows.
  • Kill switches. The agent can be stopped immediately, at any point in its workflow, by an authorized human.
  • SLA clauses. The agent's performance is contractually bound to measurable service levels — latency, accuracy, availability — with defined consequences for failure.

Humans remain in the loop at four critical junctures: approval (certifying the agent's output before it takes effect), escalation (receiving exceptions the agent cannot resolve), override (reversing an agent action when circumstances change), and post-action audit (reviewing the agent's work after the fact for accuracy and compliance).

Boards should retire the vague question "Do we have an AI strategy?" and replace it with specific questions: Who has permissions? What is auditable? Where is model provenance tracked? What are the escalation paths? Where are the kill switches?

The urgency of this framing is not theoretical. In early April 2026, roughly 512,000 lines of internal Claude Code were exposed through a misconfigured npm package. As documented by Trend Micro, the leaked code was subsequently weaponized in malware lures targeting developers — threat actors packaged convincing "leak" repositories to distribute infostealers through trusted platforms. The exposure did not just create a competitive-intelligence risk for Anthropic. It created a supply-chain attack vector that exploited the trust signals of the open-source ecosystem.

This is what weak provenance and loose control over agentic infrastructure look like at board level. If an agent touches regulated workflows — financial transactions, nuclear licensing documents, defense intelligence — and the code behind it is not governed with the same rigor as the data it processes, the institution is carrying a risk it may not be measuring.

The control layer is not a compliance cost to minimize. It is the feature that makes institutional buyers say yes.


A conservative way to bound the trillion-dollar claim

Now the market-sizing question. I have deliberately delayed this section until after the evidence, because the claim only holds if the proof points are credible and the methodology is transparent.

What the claim is: The economic value of friction removed across governed institutional workflows — backlog clearance, cycle-time compression, cost-of-delay reduction, and compliance-cost avoidance — is conservatively bounded in the hundreds of billions today and has a credible path to one trillion dollars in annual value within this decade.

What the claim is not: This is not a near-term software-revenue forecast. It is not a hardware or infrastructure TAM. It is not a projection of what AI vendors will bill. It is an estimate of the economic value that institutions will capture by embedding agents in high-friction workflows. Software vendors will capture a subset of this value pool through pricing; the rest accrues to the institutions themselves as operating leverage.

The governing metric is cost of delay per workflow-day saved. In each archetype, the primary economic value is not labor displacement but time compression. A financial close that finishes two days earlier releases working capital. A licensing submission that is ready weeks sooner reduces project carrying costs. A decision cycle that compresses from days to hours changes operational tempo. Labor savings and compliance-cost reductions are secondary levers, not the primary driver.

Three conservative assumptions:

  1. Scope limited to large, high-friction institutions. The addressable market is not "all businesses that could use AI." It is Fortune 500 and equivalent-scale enterprises, federal agencies, regulated utilities, and defense organizations — institutions where workflow friction is structural and the cost of delay is measurable in the millions per day.
  2. Workflow scope limited to auditable processes. Only workflows with explicit audit trails, regulatory requirements, or formal approval gates are included. This excludes general productivity, creative work, customer service, and other domains where value is real but harder to bound.
  3. Partial capture, not full automation. The model assumes that agents handle 30–50% of task volume in target workflows within five years, with human oversight retained at all approval and escalation points. This is deliberately below the theoretical ceiling.

Sensitivity bands:

Scenario Institutional scope Avg. workflow-days saved per year Avg. cost per delay-day Annual economic value
Low 2,000 institutions 50 days $500,000 $50 billion
Base 5,000 institutions 100 days $750,000 $375 billion
High 10,000 institutions 150 days $1,000,000 $1.5 trillion

The base case yields $375 billion in annual economic value. The high case exceeds one trillion. Even the low case — which assumes only 2,000 institutions saving 50 workflow-days per year at a conservative $500,000 per day of delay — produces $50 billion in annual value, which would make this one of the largest enterprise-value pools in technology.

Software capture will be a fraction of this. If vendors capture 10–20% of the value pool through licensing and services, the implied software TAM in the base case is $37–75 billion. That is large enough to sustain multiple platform-scale companies without requiring the headline number to be true.

The trillion-dollar claim, stated honestly: the total economic friction pool across governed institutional workflows is plausibly in this range. Software revenue will be a subset. The exact capture rate depends on pricing power, competitive dynamics, and the pace at which institutions build versus buy. But the friction is real, it is measurable, and it is being addressed by named deployments today.


Objections, counter-evidence, and adoption limits

A market thesis this large invites legitimate objections. Here are the ones I take seriously.

"These are just three examples." They are. But they are three examples across three different sectors, three different regulatory regimes, and three different institutional types — and they share a common operating pattern. The question is not whether every institution will adopt this pattern. It is whether the pattern is replicable. The structural similarities — document density, approval gates, systems-of-record dependence, and high cost of delay — are present in thousands of institutional workflows that have not yet been addressed.

"Agents will make expensive mistakes." They will. The question is whether the control layer is adequate to catch, contain, and remediate mistakes before they become material. Every case study cited here includes human-in-the-loop controls at critical junctures. The risk is not that agents are imperfect. The risk is that institutions deploy agents without adequate governance — and the Claude Code incident demonstrates what happens when control is treated as optional.

"Incumbents will resist." Some will. But the Dell case shows what happens when a CFO decides to embed agents inside finance rather than waiting for a vendor to sell them a solution. The adoption pattern that matters is not vendor-led. It is operator-led — institutions building their own capability inside their own operating model. Resistance is strongest when AI is an outside imposition. It is weakest when the function that owns the workflow also owns the agents.

"The cost-of-delay metric is too abstract." For some institutions, it is highly concrete. A nuclear energy company knows exactly what its project costs per day during a licensing delay. A CFO knows the working-capital impact of each day added to the financial close. A combatant commander knows the operational cost of decision latency. The metric is abstract only in institutions that have not yet measured their own friction — which is itself a diagnostic signal.

"Full automation will replace governance." This is the objection I reject most firmly. The history of institutional technology adoption shows the opposite: as systems become more capable, governance requirements increase, not decrease. Financial systems did not become less audited when they were digitized. They became more audited. The same dynamic will hold for agentic workflows. Governance scales with capability.


What boards should ask now

If this thesis is directionally correct — if the first trillion-dollar AI market is governed throughput, not conversational interfaces — then boards and executive teams need a different set of questions than the ones most are currently asking.

Stop asking: "Where should we pilot AI?" Start asking: "Where are approval cycles, audit trails, and document conversion throttling our operating velocity — and what would it be worth to compress them?"

Stop asking: "Do we have an AI strategy?" Start asking: "Do we have a governance framework that would allow an agent to touch our financial close, our regulatory submissions, or our supply chain — and survive audit?"

Stop asking: "Which AI vendor should we choose?" Start asking: "Should we build internal capability inside the functions that own the workflows, the way Dell built data science inside finance?"

Board readiness checklist

  • Have we identified the three to five highest-friction workflows in our operating model and quantified the cost of delay for each?
  • Do we have a governance framework that includes role-based access, segregation of duties, immutable logs, provenance tracking, exception routing, model change control, and kill switches?
  • Are our AI investments funded inside core operating budgets, or are they still sitting in innovation or pilot funds?
  • Is there a named human owner for every workflow where an agent operates — someone accountable for the agent's output, not just its deployment?
  • Have we defined the approval, escalation, override, and post-action audit points for every agent-touched workflow?
  • Are we measuring agent performance in the KPIs the business already tracks — days-to-close, submission-cycle time, exceptions per analyst, decision latency — rather than in AI-specific metrics that the board cannot benchmark?
  • Have we assessed our exposure to supply-chain risks in our agentic infrastructure — code provenance, model versioning, dependency security?
  • Can our current audit function review agent-generated work products with the same rigor it applies to human-generated ones?

This is not a capability migration that can be managed as a side project. Leaders who explain AI restructuring as simple productivity are under-describing the real shift. The most consequential change is not that existing workflows get faster. It is that the roles, skills, and organizational structures around those workflows must be redesigned — and some of the highest-performing legacy roles will still be the wrong ones for what comes next.

The institutions that move first will not be the ones with the best demos. They will be the ones that formalized first: that gave their agents a budget line, an owner, an audit trail, and a control layer that can survive scrutiny.

Formalization is the moat.


Appendix: Market-sizing assumptions and methodology

Unit of analysis: Annual economic value of friction removed from governed institutional workflows, measured as cost-of-delay reduction, with labor savings and compliance-cost avoidance as secondary components.

Institutional scope:

  • Low: 2,000 large institutions (Fortune 500, federal agencies, regulated utilities, NATO-equivalent defense organizations)
  • Base: 5,000 institutions (expanding to include large health systems, financial-services firms, and multilateral organizations)
  • High: 10,000 institutions (including mid-cap enterprises with regulated operations)

Workflow-days saved per institution per year:

  • Low: 50 (assumes 2–3 high-friction workflows with modest compression)
  • Base: 100 (assumes 4–6 workflows with meaningful compression)
  • High: 150 (assumes broad deployment across the operating model)

Cost per delay-day:

  • Low: $500,000 (conservative average across enterprise finance, regulatory, and operational workflows)
  • Base: $750,000 (weighted toward institutions with higher carrying costs)
  • High: $1,000,000 (includes defense and energy institutions with high per-day project costs)

Implied software capture: 10–20% of the economic-value pool, yielding a software TAM of $5–300 billion depending on scenario and capture rate.

Key exclusions: Consumer AI, general productivity tools, creative and marketing workflows, customer-service automation, hardware and infrastructure revenue, and model-training costs. This analysis measures only the economic value of throughput improvement in governed, auditable institutional workflows.

Validation anchors:

  • Dell's finance transformation supports the reconciliation archetype at Fortune 50 scale.
  • DOE's one-day vs. four-to-six-week document conversion supports the licensing archetype with a named, measurable compression ratio.
  • Maven's program-of-record designation at $13 billion in investment supports the defense-workflow archetype with confirmed multi-year budget commitment.

These are not projections. They are bounding assumptions. The actual market will be shaped by adoption velocity, governance maturity, vendor pricing dynamics, and institutional willingness to redesign operating models. But the friction is real. The proof points are named. And the money is already moving.


Sources and further reading:

  1. "Dell's CFO is using AI agents to run his finance team — and has helped the AI business go from $0 to $25 billion," Fortune, March 2026.
  2. "Pentagon to adopt Palantir AI as core US military system, memo says," Reuters, March 20, 2026.
  3. "US uses AI to speed up nuclear technology licensing applications," Interesting Engineering, March 2026.
  4. "Weaponizing Trust Signals: Claude Code Lures and GitHub Release Payloads," Trend Micro, April 2026.
  5. "Pentagon formalizes Palantir's Maven AI as a core military system with multi-year funding," ITPro, March 2026.
  6. "The case for boring AI," Quartz, March 2026.