China's tech giants ration AI tokens as costs bite in 2026

Share:
Audio Loading voice…
China's tech giants ration AI tokens as costs bite in 2026

Synopsis

China's top tech firms — ByteDance, Alibaba, Tencent, and Baidu — are rationing employee AI tokens and introducing co-pay models, exposing the uncomfortable reality that enterprise generative AI costs have spiralled well beyond what even the sector's biggest players can absorb without guardrails.

Key Takeaways

ByteDance now requires employees using closed-source AI models to co-pay, with the company reimbursing roughly half of external tool costs.
Annual reimbursement caps are set at approximately US$1,000 for technical staff and interns, and US$300 for non-technical employees at ByteDance .
China 's leading internet companies — including Alibaba Group Holding , Tencent Holdings , and Baidu — are rolling back previously generous AI token allowances, according to staff.
Token quotas mark a sharp reversal from the earlier culture in which high AI token consumption was treated as a productivity signal by management.
The shift reflects a broader global tension between inference costs at scale and demonstrable enterprise return on investment, a dynamic also tracked by analysts at Goldman Sachs .
Platforms such as QoderWork have emerged to help enterprises monitor and cap AI token usage, indicating the cost-control challenge extends well beyond China .
China's leading internet companies, including ByteDance, Tencent Holdings, Alibaba Group Holding, and Baidu, have begun imposing strict token quotas on employees as surging generative AI costs force a rethink of the freewheeling usage culture that defined the sector's early AI adoption phase, according to staff who spoke on condition of anonymity.

From badge of honour to rationed resource

When generative AI first swept through China's technology sector, workers faced a clear directive from management: use AI and use it often. Consuming vast numbers of AI tokens — the basic unit of computing power needed to process text or write code — became a badge of honour, with companies often viewing a high token count as a sign that employees were productive and taking initiative. That era is now over. Computing allowances are being sharply rolled back as firms introduce formal token quotas to rein in rising AI-related expenditure, according to multiple employees who described the shift.

ByteDance leads with co-pay model

At ByteDance, one employee said staff using closed-source models must now co-pay out of pocket. The company reportedly reimburses roughly half of external tool costs, capping annual payouts at around US$1,000 for technical staff and interns, and US$300 for non-technical employees, according to the staff member. The co-pay structure effectively transfers a portion of the cost burden directly to workers — a notable reversal from the earlier posture in which unlimited or near-unlimited AI access was positioned as a workplace benefit.

Why it matters

The shift signals that the economics of generative AI at enterprise scale are proving far harder to absorb than initially anticipated. Token costs accumulate rapidly when thousands of engineers and non-technical staff run large language model queries throughout the workday, and the aggregate bill can rival or exceed traditional software licensing expenses. The pullback also complicates the narrative that China's tech giants have been aggressively out-deploying their Western counterparts on internal AI adoption. Rationing suggests that raw usage volume was never a sustainable proxy for productivity or return on investment.

The competitive backdrop

The token-rationing trend arrives as global firms, including those tracked by Goldman Sachs in its enterprise AI spending analyses, grapple with the same fundamental tension: the cost of inference at scale versus the productivity gains that justify it. Platforms such as QoderWork have emerged specifically to help enterprises monitor and cap AI token consumption, suggesting the problem is industry-wide. For Alibaba, Tencent, Baidu, and ByteDance, the stakes are particularly acute because each also operates its own foundation models and cloud infrastructure — meaning internal token consumption directly competes with revenue-generating external workloads for the same compute capacity.

What's next

The most immediate question is whether token quotas will dampen the pace of internal AI tool adoption just as these companies are racing to embed generative capabilities into core products. Employees subject to hard caps may self-censor usage, slowing the feedback loops that help enterprises identify the highest-value AI applications. Watch for whether ByteDance's co-pay model spreads to peers, and whether any company publicly quantifies the cost savings achieved — a disclosure that would offer the clearest window yet into how expensive enterprise-scale AI deployment has become in China.

Point of View

And the productivity gains that were supposed to justify open-ended AI spending remain difficult to quantify at the individual employee level. What mainstream coverage tends to underplay is that companies like ByteDance and Alibaba are simultaneously the consumers and the providers of foundation model compute — meaning internal token rationing is also a capacity-allocation decision that directly affects their cloud revenue potential. The co-pay model pioneered by ByteDance is particularly telling: it outsources cost discipline to individuals rather than building smarter usage policies, which risks chilling exactly the experimental behaviour that generates the most valuable enterprise AI use cases. If this trend hardens, it could structurally slow the feedback loops that Chinese tech giants need to stay competitive with US hyperscalers on model fine-tuning and application development.
NationPress
16 Sept 2026

Frequently Asked Questions

Why are Chinese tech companies introducing AI token quotas?
China's leading internet companies are introducing AI token quotas because the aggregate cost of generative AI usage by thousands of employees has grown unsustainably large. What began as an encouraged behaviour — high token consumption as a proxy for productivity — has become a significant line item that firms including ByteDance , Alibaba , Tencent , and Baidu are now actively seeking to control.
How does ByteDance's AI reimbursement policy work?
ByteDance reimburses roughly half of what employees spend on external closed-source AI tools, according to a staff member who spoke anonymously. Annual reimbursement is capped at around US$1,000 for technical staff and interns, and US$300 for non-technical employees, effectively requiring workers to co-pay for any usage above those thresholds.
What is an AI token and why does it cost money?
An AI token is the basic unit of computing power that large language models use to process text or generate code — roughly corresponding to a fragment of a word. Every query to a generative AI model consumes tokens, and at enterprise scale, with thousands of employees querying models throughout the day, the cumulative compute cost can be substantial.
Does AI token rationing affect employee productivity?
Hard token caps risk dampening internal AI adoption by encouraging employees to self-censor usage to avoid out-of-pocket costs. This could slow the feedback loops companies rely on to identify the highest-value AI applications, potentially undermining the productivity gains that were the original justification for enterprise AI investment.
Is AI cost rationing a problem only in China?
No — the challenge is global. Analysts at Goldman Sachs have tracked the tension between enterprise AI spending and demonstrable returns across markets, and platforms such as QoderWork have emerged specifically to help companies worldwide monitor and cap token consumption. China 's tech giants are among the first to implement formal co-pay structures, but the underlying cost pressure is industry-wide.
Nation Press
The Trail

Connected Dots

Tracing the thread behind this story — newest first.

8 Dots
  1. Latest 4 weeks ago
  2. 1 month ago
  3. 1 month ago
  4. 2 months ago
  5. 2 months ago
  6. 2 months ago
  7. 2 months ago
  8. 4 months ago
Google Prefer NP
On Google