China's NDA unveils 2028 plan to build AI training data ecosystem

Share:
Audio Loading voice…
China's NDA unveils 2028 plan to build AI training data ecosystem

Synopsis

China's National Data Administration has unveiled a national plan to build validated AI training data sets across nine industries by 2028 — including embodied AI and autonomous driving — a direct state-level response to the global data drought threatening next-generation AI development.

Key Takeaways

China's National Data Administration published a draft nationwide AI data plan on Monday, 9 June 2026 .
The road map targets a validated ecosystem of industry-specific data sets to be operational by 2028 .
Covered sectors include scientific research, manufacturing, agriculture, energy, transport, finance, healthcare, education and e-commerce .
Cutting-edge domains such as embodied AI, autonomous driving, low-altitude aviation and biomanufacturing are explicitly included.
The plan calls for multimodal data — text, code, images, audio and video — to support complex reasoning, agentic AI and intelligent robotics.
The initiative is framed as a pillar of Beijing's AI Plus strategy , using industrial data as a structural advantage amid global chip export controls on China .

China's National Data Administration (NDA) has released a sweeping draft plan to dramatically expand the country's supply of high-quality AI training data, positioning data as a core strategic asset in the global artificial intelligence race. Published on Monday, 9 June 2026, the initiative directly addresses a looming global shortage of data needed to power next-generation AI models.

What the plan covers

The draft road map, published online by the National Data Administration, outlines a framework for expanding the supply, circulation and commercialisation of industry-specific data sets. The initiative is designed to anchor Beijing's AI Plus strategy — a top-down mandate to integrate artificial intelligence into the industrial fabric of the world's second-largest economy.

The plan calls for a broad expansion into multimodal data — spanning text, code, images, audio and video — to train advanced systems capable of complex reasoning, agentic behaviour and controlling intelligent robots.

Target sectors by 2028

By 2028, the agency aims to field a validated ecosystem of data sets covering bedrock sectors including scientific research, manufacturing, agriculture, energy, transport, finance, healthcare, education and e-commerce. Cutting-edge frontiers such as embodied AI, autonomous driving, low-altitude aviation and biomanufacturing are also explicitly included in the scope.

The breadth of sectors targeted signals that Beijing intends to use proprietary industrial data — accumulated across its vast manufacturing and services economy — as a structural competitive advantage against Western AI developers.

Why it matters

AI developers globally are confronting what analysts describe as a data drought: the stock of high-quality, publicly available text and image data used to train large language models is reportedly nearing exhaustion. China's move to institutionalise data collection and commercialisation at a national level could give its AI developers a sustained supply advantage, particularly in industrial and scientific domains where proprietary data is scarce worldwide.

The plan's emphasis on multimodal and embodied AI data is especially significant, as robotics and autonomous systems are widely seen as the next major frontier in AI commercialisation.

The competitive backdrop

The announcement comes as Chinese AI firms including DeepSeek, Baidu and Alibaba compete aggressively with US counterparts such as OpenAI, Google and Anthropic. Access to differentiated training data — rather than raw compute — is increasingly seen as the decisive variable in the next phase of AI development, particularly as chip export controls constrain China's access to leading-edge semiconductors.

What's next

The plan is currently in draft form, and the National Data Administration is expected to invite public comment before finalising the road map. Stakeholders in sectors from autonomous vehicles to precision agriculture will be watching closely as Beijing moves to operationalise what could become the world's largest state-coordinated AI data infrastructure. The pace of implementation — and whether private-sector data holders comply — will determine how quickly China can convert this policy ambition into a tangible model-training advantage.

Point of View

A state-orchestrated data advantage in industrial and scientific domains could partially offset that constraint. What mainstream coverage underweights is the embodied AI and robotics angle — proprietary sensor, motion and environmental data from China's vast manufacturing base could prove more strategically durable than any single model benchmark. The plan also signals that Beijing views data commercialisation, not just data collection, as essential — meaning domestic AI firms may gain preferential access to data markets that foreign competitors cannot enter. The real test will be whether state-held data actually flows to private developers efficiently, or whether bureaucratic friction neutralises the policy's ambition.
NationPress
26 Jul 2026

Frequently Asked Questions

What is China's National Data Administration AI plan?
China's National Data Administration released a draft road map on 9 June 2026 to expand the supply, circulation and commercialisation of industry-specific AI training data sets, targeting a validated ecosystem across nine major sectors by 2028 . The plan is a key pillar of Beijing's AI Plus strategy .
Why is China building its own AI training data supply?
AI developers worldwide are facing a looming shortage of high-quality training data as publicly available text and image sources near exhaustion. China is institutionalising data collection at the national level to give its AI firms a sustained supply advantage, particularly in industrial and scientific domains.
Which sectors are covered by China's AI data plan?
The plan covers scientific research, manufacturing, agriculture, energy, transport, finance, healthcare, education and e-commerce , as well as cutting-edge frontiers including embodied AI, autonomous driving, low-altitude aviation and biomanufacturing .
What is multimodal AI data and why does China's plan include it?
Multimodal data spans text, code, images, audio and video, and is used to train AI systems capable of complex reasoning, agentic behaviour and controlling intelligent robots. China's plan prioritises multimodal data to support next-generation AI including humanoid robotics and autonomous vehicles.
How does China's AI data plan affect global AI competition?
By building proprietary industrial data sets that Western developers cannot easily replicate, China aims to create a structural competitive advantage even as chip export controls limit its access to leading-edge semiconductors. The initiative directly challenges the assumption that compute alone determines AI leadership.
Nation Press
The Trail

Connected Dots

Tracing the thread behind this story — newest first.

8 Dots
  1. Latest 3 days ago
  2. 1 week ago
  3. 1 week ago
  4. 3 weeks ago
  5. 3 weeks ago
  6. 3 weeks ago
  7. 1 month ago
  8. 3 months ago
Google Prefer NP
On Google