Token Cost Is the Real Infrastructure. The Rest Is Real Estate.
Most boards think AI infrastructure means data centers and GPUs. But the real gate is the unit economics of the agentic execution loop: tokens consumed per outcome, retries, tool calls, and re-plans. Caching turns tokens into margin, and digital trade agreements can function like a cost discount fo…

The most expensive thing about your AI strategy is probably not on the cap-ex line. It is the token bill you have not modeled yet.
Aaron Levie, the CEO of Box, said this week at Semafor's Tech summit that AI token costs arrived "overnight," catching companies that had already made coding agents indispensable. Box, he said, would "cease to function" without them. Cisco's Jeetu Patel, on the same stage, noted that an AI agent doing a knowledge-worker task can consume close to five times the bandwidth of the human it replaced. Take those two numbers seriously and the standard infrastructure conversation starts to look misdirected.
My read is that most boards still think AI infrastructure means data centers, GPUs, and power contracts. That is real estate. The infrastructure that actually decides whether you can run an agentic workflow at margin is the unit economics of the loop itself: tokens consumed per outcome produced, multiplied by every retry, every tool call, every re-plan. CFOs are quietly discovering that AI agents, unlike junior analysts, file their expense reports in real time.
This matters because the cost is now a design object, not a fate.
Caching turned tokens into a margin lever
Look at the OpenAI rate card. The headline numbers per million input and output tokens are not the interesting part. The interesting part is prompt caching, which on most current models cuts the cost of cached input tokens by roughly half. Reasoning models meter differently again. The card is no longer a price list. It is a manual for designing margin into your workflow.
That changes who owns the token bill. If you write your agent loops naively, with a fresh context window on every call and no memory of what was already paid for, your unit economics decay with every iteration. If you architect with cache hits, deterministic tool routing, and tight prompts, the same workflow can cost less every quarter as the platform's caching and quantization improve underneath you.
This is the layer I keep coming back to. The competitive edge in agentic systems is not whether you have agents. It is whether your agents have been designed to be cheap. Marketing teams running content engines, CX teams running deflection agents, RevOps teams running outbound orchestration are all about to discover that two organizations with the same use case can have a 5x to 10x spread in tokens per outcome. One scales. The other quietly retreats.
NVIDIA's research has been blunt about this for a while. Cost per token is the metric that matters for AI factories, because everything downstream, from how often you can retrain to how long a reasoning chain you can afford, is a derivative of it. Even SpaceX's prospectus filings describing the AI-adjacent business case treat token cost as the constraint that defines what is buildable. Capital markets are catching up to what operators are already paying.
Where geopolitics enters the workflow bill
Here is where most analysis stops, with a tidy lesson about caching strategy. The more interesting move is to follow the cost curve out of the data center and into the trade office.
South Korea's president, Lee Jae-myung, is currently on a European tour that takes him through EU leadership in Brussels, the Vatican, and the G7 meeting in France. The political narrative is restoration after his predecessor's martial-law impeachment, a country signaling a more confident global posture. That is real, but it is not what builders should be reading.
The agenda items are trade, technology, and defense. Underneath those headers sits the EU-South Korea Digital Trade Agreement, being negotiated and shaped in exactly this kind of summit corridor. Its provisions on cross-border data flows, source code protection, and digital service treatment are not abstract legal furniture. They are the rules that will decide whether a Korean enterprise can run inference against a European customer dataset without paying a tax in latency, legal review, and re-architecture every time the data crosses a border.
Every one of those frictions shows up as tokens. More retries. More routing. More duplicated context. More compliance text injected into prompts. More hedging chains because the agent does not trust its own data lineage. A trade agreement that smooths data flow between two AI-heavy economies is, functionally, a discount on agentic workflow cost.
That is the reframe worth holding. The diplomatic itinerary of a semiconductor and AI powerhouse is not a story about charm. It is a story about which jurisdictions get to host the cheapest, most predictable agentic workloads over the next five years. Trade agreements are becoming cost documents. The people writing them probably do not know that yet. The people running the workloads should.
What this changes for marketing and CX teams
If you run an AI-native marketing or customer experience function, the implication is concrete.
Stop benchmarking your agentic content engine by volume or model choice. Benchmark it by tokens per converted outcome, tokens per resolved ticket, tokens per qualified handoff. That is your true unit cost, and it is the only number that scales linearly with your ambition. Teams that publish a thousand AI-generated assets a week without knowing their cost per booked meeting are not running an agentic program. They are running a budget leak with good UX.
Then look at where those workloads physically run, and under which legal regime. If your CX agent serves European customers from a US inference endpoint, every hedged response and every duplicated context window is a tax you pay on regulatory uncertainty. If a digital trade agreement between your home market and your customer's market reduces that uncertainty, your per-outcome cost falls without you touching a single prompt. That is real infrastructure, and it does not appear on any architecture diagram.
This is also where the integration question matters. Earlier I argued that marketing's integration debt is becoming a product category and that compute geography is becoming the routing layer for agentic workloads. Token cost is the layer that sits on top of both. Geography tells you where the workload can legally run. Integration tells you whether the loop is governed. Token economics tell you whether the whole thing pays for itself.
The discipline of knowing what you are paying for
There is a quiet operating discipline underneath all of this, and it has nothing to do with caution.
Speed is the right posture. Building fast with frontier models is the right posture. The mistake is letting speed disguise the bill. A leader who cannot say, in plain numbers, what a given agentic workflow costs per outcome today, and what would have to change in their prompts, their caching, their data residency, or their trade jurisdiction to cut that cost in half, is not in command of the system. They are renting confidence from it.
The next infrastructure conversation in your organization should not be about more compute. It should be about the unit economics of the workflow loop, and which agreements, technical and diplomatic, make those economics predictable enough to scale against. That is the bet worth making and the bet worth understanding.
The data centers will get built either way. The question is whose token bill stays legible long enough to turn into a business.
Sources and further reading
- How Autonomous AI Agents Become Secure by Design With NVIDIA OpenShell: historically indexed source that semantically overlaps with the current article frame.
- SpaceX Second Amended and Restated Preliminary Base Prep Prospectus: historically indexed source that semantically overlaps with the current article frame.
- Compute Geography Is Becoming the Routing Layer for Agentic Workloads: published AI Stoic archive memory supporting crosslinking and non-repetition.
- Marketing's Integration Debt Is Becoming a Product Category: published AI Stoic archive memory supporting crosslinking and non-repetition.
Reader account
Join the conversation
Sign in with a private email link to manage preferences and leave a comment.

Comments