When the Input Goes to Zero, Who Keeps the Margin?

DeepSeek just taught the market something railroads taught it a century ago: the companies that survive a commodity collapse are almost never the ones selling the commodity.

The sequence of events over the last four weeks deserves to be read as a single argument. On July 30, OpenAI cut the input token price of GPT-5.6 Luna from $1 to $0.20 per million tokens, a reduction of 80%, which it described as being made possible by improvements in model efficiency. Hours later, DeepSeek released V4-Flash and undercut that new price by a further 30%. DeepSeek charged $0.14 per million input tokens and $0.28 per million output tokens for V4-Flash, while Anthropic’s top model was priced far higher than that. The gap between the two was not a rounding error. It was a structural claim about who controls the economics of AI.

Then came the reversal. DeepSeek’s new pricing, effective August 16, took V4-Pro output tokens from $0.87 per million to $3.96 during peak hours, a 355 percent increase, with $1.98 off-peak. V4-Flash output jumped from $0.28 to $1.32 per million at peak, a 371 percent increase. The company that seemed to be racing inference costs toward zero is now reportedly planning a one-gigawatt data center in Inner Mongolia and has been discussed in the press as a potential future IPO candidate. Capital for that data center and in-house AI inference chips would reduce dependence on third-party compute. For a company approaching public markets, unit economics that a prospectus must defend look different from promotional pricing designed to capture developer adoption.

This is where the railroad analogy earns its keep. When steel rail prices collapsed in the 1880s, it did not destroy railroad profits. It destroyed the margins of steelmakers. The railroads, as operators of a network that users could not route around, held their pricing power. The input commoditized. The routing did not.

The AI stack is splitting the same way. Anthropic has publicly said its run-rate revenue crossed $47 billion earlier this year. Separately, analysts and industry research have pointed to inference margins climbing from roughly the high 30s to above 70% over the past year as enterprise usage surged. That is a business growing into the commodity storm, not shrinking from it. The reason is that Anthropic does not primarily sell tokens. It sells the capability ceiling, safety properties, and the workflow integrations that enterprises are unwilling to rebuild around a cheaper Chinese substitute.

The harder question concerns Nvidia and the hyperscalers. The four major hyperscalers have guided to roughly $725 billion of capital expenditures in 2026, up about 77 percent from roughly $410 billion in 2025, according to multiple third-party compilations of company guidance. Capital expenditures are growing far faster than operating cash flow, squeezing aggregate free cash flow. The bull case for that spending is that token volume will expand faster than price declines. OpenRouter data reported V4-Flash first for the week of July 27 to August 2 at 7.22 trillion tokens, and OpenCode has reported daily volumes above 8 trillion tokens in early August. Volume is not the problem. Whether the revenue that volume generates justifies the infrastructure being built to serve it is.

Nvidia reports fiscal Q2 2027 results today, August 26, 2026. Consensus expectations vary by data source but cluster around roughly $92 billion in revenue and about $2.08 in earnings per share. The guide for Q3 will matter more than the beat. What every investor is really reading is whether hyperscaler customers are pulling forward or pushing back. DeepSeek’s price hike removes one argument for pulling back. Its capacity plans add a new argument for why Nvidia’s backlog holds.

The investors who built lasting wealth from the bandwidth collapse of the early 2000s were not the ones who shorted fiber. They were the ones who identified which layer of the network retained pricing power when transit costs went to zero. Applications and operating systems, not raw pipe. The AI version of that question is now live. Tokens are the pipe. The layer that keeps the economics is the one that owns the workflow, the brand, or the proprietary data the model is trained on. Companies that have those things will compound. Companies that sell commoditized compute to companies that have those things will remain captive to the capex cycle. And companies sitting in the middle, application builders whose only moat was access to cheap tokens, face the hardest path of all.