China unveils corpus to push advanced tech research
By LI MENGHAN | China Daily | Updated: 2026-07-21 10:27
China, on Friday, unveiled a national scientific corpus infrastructure at the 2026 World Artificial Intelligence Conference in Shanghai, aiming to provide the high-quality data needed to support the next generation of AI-driven scientific research.
Called the Aocang S&T Corpus, the platform was jointly developed by the National Science Library of the Chinese Academy of Sciences and nearly 100 research institutions. It is designed as a national repository of high-quality scientific data covering all major disciplines, providing AI systems with reliable information for research and innovation.
The name Aocang comes from the largest state granary of the Qin Dynasty (221-206 BC), built in present-day Henan province to supply grain for the newly unified empire. Qu Jiansheng, director of the National Science Library, said the project draws inspiration from the ancient granary by serving as a strategic repository for scientific knowledge in the digital age.
The launch comes as AI for Science — the use of AI to help scientists conduct research and make discoveries more efficiently — has become a national strategic priority. China is accelerating the development of AI-enabled scientific research to support its goal of becoming a global leader in scientific innovation, Qu said.
Unlike the vast amount of information collected from the internet, scientific data is often scattered across different institutions, stored in different formats and requires specialized knowledge to interpret. These challenges make it difficult to use such data to train AI models.
Aocang is designed to solve this problem by organizing scientific information into a unified format that AI systems can more easily understand and process.
"Aocang holds irreplaceable, foundational and strategic value for strengthening the data foundation for China's AI development and advancing high-level scientific and technological self-reliance," Qu said.
Drawing on the CAS' extensive research resources, Aocang integrates more than 320 petabytes of scientific data, 150 million research papers, 120 million patents and 5 million scientific books.
The platform is organized into a national core database and a number of specialized databases covering fields ranging from basic science to industrial technology. It can process many types of scientific information, including research papers, mathematical formulas, spectral data and waveforms.
To make the information more useful for AI, developers built a three-stage processing system that cleans, organizes and enriches raw scientific data before turning them into training materials. According to Qian Li, deputy director of the National Science Library, the process enables AI models to better understand scientific concepts and relationships rather than simply recognize words or patterns.
"These knowledge-rich datasets can strengthen large AI models' expertise in specialized fields, improve their reasoning ability and help them better understand scientific disciplines," Qian said.
More than 520 proprietary software tools support the platform's automated system for data processing, quality control and secure data sharing. The developers said the corpus will continue to expand as more research institutions contribute data.
Initial tests show that Aocang has improved the reasoning capabilities of ScienceOne, a foundational AI model for scientific research developed by the CAS.
The corpus is also being integrated into several commercial AI models, including Doubao, Qwen and Ant Group's medical AI model.
During its first phase, the project aims to integrate more than 36 billion scientific resource records across 2,883 high-quality scientific corpora and 68 PB of research data, supporting more than 500 research and industry application scenarios.
Qu said the academy will continue expanding the corpus, improving its data-processing capabilities and exploring operating models that combine public access with sustainable commercial services.
"We hope the corpus will support the development of new quality productive forces and provide a reliable, world-class data foundation for scientific innovation around the world," he said.
limenghan@chinadaily.com.cn





















