Enterprise AI inference costs hit 2026 low amid global price war, DeepSeek surge
Synopsis
Key Takeaways
Enterprise AI inference costs have dropped to their lowest point of 2026, driven by an intensifying global price war and rapid adoption of low-cost Chinese open-source models, according to research by investment bank Jefferies. The findings, citing data from US-based research firm Silicon Data, signal a structural shift in how businesses budget for AI workloads.
Prices at a record low for 2026
Average inference prices — measured per million tokens, or chunks of data processed by a model — ranged between US$1.16 and US$1.18 from August 6 to 8, marking the lowest level recorded this year, Jefferies said on Monday. The decline is steep: unit costs stood at US$2.04 on May 31 and had already fallen to US$1.45 in late July before dropping further. Silicon Data's index tracks pricing across business API providers and open-weight inference platforms used by software developers.
Why it matters: the price war accelerates
The price decline coincided with an 'increasing emphasis on cost efficiencies' across both the US and Chinese tech ecosystems, Jefferies analysts led by Thomas Chong noted. OpenAI escalated the price war last month by slashing rates for its latest GPT-5.6 model series by up to 80 per cent. Meanwhile, Anthropic's Claude Opus 5 reportedly delivered performance comparable to its flagship Fable 5 model at half the price, according to the Jefferies report.
The competitive backdrop: Chinese open-source pushes costs lower
On the open-source front, Chinese firms are aggressively expanding the boundaries of affordable computing. DeepSeek, alongside peers such as Zhipu AI, Moonshot AI, and MiniMax, has contributed to a wave of low-cost, high-performance models that enterprises can deploy without paying premium API rates. Platforms tracked by Silicon Data — including OpenRouter — aggregate these options, giving developers direct price visibility and intensifying competitive pressure on incumbent providers like Google and Tencent Holdings.
What's next: further compression likely
The convergence of Western price cuts and Chinese open-source momentum suggests inference costs may compress further through the second half of 2026. Enterprises that locked in higher-cost AI contracts earlier this year face the most immediate exposure, while cloud-native developers stand to benefit from dramatically lower build costs. Analysts and industry observers will be watching whether OpenAI, Anthropic, and major cloud hyperscalers respond with another round of cuts — or shift competition to capability and reliability rather than price alone.