China's NDA unveils 2028 plan to build AI training data ecosystem
Synopsis
Key Takeaways
China's National Data Administration (NDA) has released a sweeping draft plan to dramatically expand the country's supply of high-quality AI training data, positioning data as a core strategic asset in the global artificial intelligence race. Published on Monday, 9 June 2026, the initiative directly addresses a looming global shortage of data needed to power next-generation AI models.
What the plan covers
The draft road map, published online by the National Data Administration, outlines a framework for expanding the supply, circulation and commercialisation of industry-specific data sets. The initiative is designed to anchor Beijing's AI Plus strategy — a top-down mandate to integrate artificial intelligence into the industrial fabric of the world's second-largest economy.
The plan calls for a broad expansion into multimodal data — spanning text, code, images, audio and video — to train advanced systems capable of complex reasoning, agentic behaviour and controlling intelligent robots.
Target sectors by 2028
By 2028, the agency aims to field a validated ecosystem of data sets covering bedrock sectors including scientific research, manufacturing, agriculture, energy, transport, finance, healthcare, education and e-commerce. Cutting-edge frontiers such as embodied AI, autonomous driving, low-altitude aviation and biomanufacturing are also explicitly included in the scope.
The breadth of sectors targeted signals that Beijing intends to use proprietary industrial data — accumulated across its vast manufacturing and services economy — as a structural competitive advantage against Western AI developers.
Why it matters
AI developers globally are confronting what analysts describe as a data drought: the stock of high-quality, publicly available text and image data used to train large language models is reportedly nearing exhaustion. China's move to institutionalise data collection and commercialisation at a national level could give its AI developers a sustained supply advantage, particularly in industrial and scientific domains where proprietary data is scarce worldwide.
The plan's emphasis on multimodal and embodied AI data is especially significant, as robotics and autonomous systems are widely seen as the next major frontier in AI commercialisation.
The competitive backdrop
The announcement comes as Chinese AI firms including DeepSeek, Baidu and Alibaba compete aggressively with US counterparts such as OpenAI, Google and Anthropic. Access to differentiated training data — rather than raw compute — is increasingly seen as the decisive variable in the next phase of AI development, particularly as chip export controls constrain China's access to leading-edge semiconductors.
What's next
The plan is currently in draft form, and the National Data Administration is expected to invite public comment before finalising the road map. Stakeholders in sectors from autonomous vehicles to precision agriculture will be watching closely as Beijing moves to operationalise what could become the world's largest state-coordinated AI data infrastructure. The pace of implementation — and whether private-sector data holders comply — will determine how quickly China can convert this policy ambition into a tangible model-training advantage.