Citrini
· Citrini Research
· June 08, 2026 at 13:06
· ⏱ 10 min read
| Read on Substack ↗
Summary
The AI industry is entering a 'token panic' where explosive token usage from agents and reasoning models is causing customer costs to skyrocket, leading to corporate pushback and a shift towards usage-based billing, open-source alternatives, and efficiency optimization. This means the AI trade is rotating from pure infrastructure spending toward cost-conscious winners like efficient model providers and local inference, while frontier labs face pricing pressure despite continued revenue growth.
•Uber burned through its entire 2026 AI budget in just four months, exemplifying corporate cost overruns from token consumption.
•Anthropic's ARR increased 5x since start of 2026 to $45 billion in May, but its CEO acknowledged cost became a 'huge issue' only recently.
•Deepseek V4 Pro and V4 Flash are 10×–25× cheaper than Opus 4.8 and GPT 5.5, and Deepseek overtook Anthropic on OpenRouter in tokens processed.
•Cursor released a coding model post-trained on xAI compute that is comparable to GPT-5.5 and Opus 4.7 at 10× lower cost per task, using a Chinese open-source base by Moonshot.
•OpenAI, Google, and GitHub Copilot all shifted to usage-based billing between April and June 2026, ending unlimited-consumption pricing.
•Nvidia released the Nemotron family of open-source models, including compact versions optimized for local deployment and specialized agentic uses.
The article explicitly notes Nvidia released the 'latest Nemotron family which includes advanced general-purpose models as well as highly efficient, compact versions optimized for local deployment' —
The article explicitly notes Nvidia released the 'latest Nemotron family which includes advanced general-purpose models as well as highly efficient, compact versions optimized for local deployment' — aligning with the emerging theme of cost-efficient, open-source inference that could expand Nvidia's software ecosystem without cannibalizing its hardware dominance.
Risk: Open-source models could reduce demand for high-end Nvidia GPUs if customers opt for cheaper inference, but Nvidia's hardware remains the backbone for training and most frontier inference.