The workload has already moved.
On OpenRouter, the routing layer that sits between apps and dozens of AI providers, Chinese-origin models now account for roughly 46.4% of routed tokens against about 35.7% for US-origin ones, per OpenRouter’s public rankings. DeepSeek alone routes 17.6% of the platform’s traffic, near 5.13 trillion tokens a week, making it the single largest vendor. Alibaba’s Qwen adds another 13.9%. A year ago American models held roughly 70% of the same pipe. A CNBC investigation in July put the Chinese weekly peak at 46%, up from an 11% trailing-year average.
This is not a benchmark result. It is a routing decision, made millions of times a day, by people paying the bill.
Now look at the money.
Vercel’s AI Gateway Production Index, which measures live traffic from deployed applications rather than test harnesses, tells the other half. Open-weight models ran 29% of gateway tokens in June, up from 11% in April. They accounted for under 4% of dollar spend. Nearly a third of the work, one twenty-fifth of the revenue.
The reason is price. DeepSeek V4 Flash costs about $0.14 per million input tokens. OpenAI’s GPT-5.5 runs $5.00 for the same. Chinese open-weight models are consistently 60% to 90% cheaper than frontier American offerings. When a model is that cheap, high-volume, low-stakes work floods to it: summarization, classification, draft generation, internal tooling. The token counter spins. The invoice barely moves.
Where the value actually sits.
The production data is unusually clear about where frontier labs still win. On Vercel’s gateway, Anthropic captured 61% of spend on 32% of tokens, and more than 72% of spend in every high-stakes category: coding agents, back-office agents, app generation. These are the jobs where a wrong answer is expensive, where the buyer will pay thirty times per token to shave the error rate. Value capture is concentrating exactly where mistakes cost money.
So “China is catching up” misreads the board. On raw capability the gap has narrowed, but that is not what the two curves are measuring. Token share measures where inference is commoditizing. Dollar share measures where it is not. Both are true at once, and they are diverging on purpose.
The base is widening, not just the leaders. The Ramp AI Index shows the share of AI-spending businesses paying for model-serving platforms, the layer that hosts open-weight and Chinese models, rose to 6.1% in July from 4.5% in January. The cheap tier is broadening across the mid-market, not confined to a few cost-obsessed startups.
Which curve you underwrite.
For a finance audience the question is not who has the best model. It is which curve predicts the business. Two readings.
Bet on tokens, and inference is commoditizing faster than the frontier labs’ revenue admits. Open weights set a price ceiling. Every quarter DeepSeek and Qwen absorb more volume, the frontier premium has to be justified by a shrinking set of tasks. Gross margins on undifferentiated inference trend toward the cost of electricity, and platform revenue, not model revenue, becomes the prize.
Bet on dollars, and the frontier moat is holding where it counts. The high-stakes work is sticky, the switching cost is real, and 96% of the money still flows to models that charge for reliability. Volume share is a vanity metric until it converts into work that pays.
The honest position: both are right for now, and the tension resolves task by task. Every workload that migrates from “expensive and careful” to “cheap and good enough” moves revenue from the dollar curve to the token curve. The frontier labs are betting they can invent new expensive work faster than their old expensive work commoditizes. That is the entire wager.
The one variable that overrides price.
Policy can cap the trend regardless of economics. The FY2026 NDAA bars the Defense Department and its contractors from using DeepSeek or its parent High-Flyer. Pending bills would keep DeepSeek off federal devices and force the Federal Acquisition Security Council to publish a list of models from adversarial nations, with agencies needing an exemption to touch them. Texas, New York, Virginia, and several other states have already banned DeepSeek from government networks. The export-control picture runs both directions, but the pressure on Chinese weights is the active front.
None of this touches a private company running Qwen behind its own firewall today. There is no nationwide corporate ban, and open weights are hard to ban by design: once the file is downloaded, there is no server to shut off. But it fences the fastest-growing tier out of government, defense, and regulated procurement, precisely the high-stakes, high-margin work the American labs already dominate. The policy overhang does not reverse the scissors. It widens them.
The workload has gone to China. The wallet has not. For the next year and a half, the AI business is the story of how fast those two lines converge, and how much money is trapped in the gap while they do.
Discussion
Sign in to join the discussion.
No comments yet. Be the first to share your thoughts.