DeepSeek V4 triggers 99% API price cuts, reshaping China AI market
Synopsis
Key Takeaways
DeepSeek's ultra-low-cost V4 models are forcing a sweeping repricing across China's artificial intelligence industry, compelling rivals to slash fees and experiment with new monetisation models as competition in the domestic market intensifies. The latest wave of cuts, reported on 8 June 2026, signals that the AI pricing war ignited by DeepSeek shows no sign of abating.
Xiaomi leads the latest price retreat
Xiaomi, the Shenzhen-based smartphone and electric-vehicle maker, has emerged as one of the most prominent responders, cutting application programming interface (API) costs for its MiMo-V2.5 model by as much as 99 per cent from previous levels. The steep reduction immediately translated into surging developer adoption, with MiMo-V2.5 climbing to sixth place on US-based model marketplace OpenRouter within days of the announcement.
The model processed 1.7 trillion tokens in the seven days to Monday, representing growth of more than 999 per cent from the prior week — a figure that underscores how sharply price sensitivity shapes developer behaviour in the current environment. The premium variant, MiMo-V2.5-Pro, also recorded a notable uptick in usage over the same period.
Why it matters
The cascade of price reductions illustrates how DeepSeek's cost-efficient architecture has reset baseline expectations for what AI inference should cost, putting sustained pressure on every Chinese model provider's unit economics. Cloud platforms that built revenue forecasts around higher per-token margins are now navigating a structurally lower-price environment with no clear floor in sight.
For enterprise customers and independent developers, the dynamic creates an unprecedented window of cheap access to frontier-class models — but it raises questions about the long-term viability of providers who cannot offset margin compression with scale or differentiated services.
MiniMax bets on subscriptions alongside token billing
Not every player is competing purely on price. AI unicorn MiniMax on Monday launched its next-generation flagship model, MiniMax M3, pairing conventional token-based billing with subscription plans ranging from US$7.24 to US$69.28 per month. The dual-track approach reflects a broader industry search for monetisation models that can survive in a market where raw inference costs are collapsing.
The move positions MiniMax as a test case for whether predictable subscription revenue can provide a more stable foundation than usage-based pricing alone, a question that will be closely watched by investors and competitors alike.
The competitive backdrop
The repricing wave is unfolding against a backdrop of intense rivalry among China's AI developers, many of whom are racing to secure developer mindshare before the market consolidates. DeepSeek's V4 has effectively become the industry's cost benchmark, forcing even well-capitalised players to respond or risk losing relevance on third-party marketplaces and enterprise procurement shortlists.
The transition has not been without friction, according to reports, as providers balance the need to attract volume against the imperative to sustain the infrastructure investment required to remain competitive.
What's next
The central question now is whether the current pricing floor holds or whether further cuts — potentially driven by the next generation of efficiency improvements — push margins even lower. Investors and cloud providers alike will be watching whether subscription-hybrid models like MiniMax M3's can gain enough traction to offer a credible alternative to the race to zero on per-token costs.