Chinese automakers are increasingly competing over how efficiently they can identify and use valuable autonomous-driving data, as the industry moves beyond simply collecting larger volumes of road footage.
An Level 4 test vehicle can generate about 20 terabytes of raw data each day, but most of it consists of routine driving, such as highway cruising and straight-road travel. The most important material is often found in rare “edge cases,” including a pedestrian suddenly emerging from behind an obstruction or an unexpected lane change.
China’s intelligent-driving industry is therefore shifting from manual, trigger-based data mining toward semantic retrieval, active learning and world-model-driven training. Industry data presented at the 2026 World Intelligent Connected Vehicles Conference showed that Level 2 intelligent-driving systems had reached a 70.5% penetration rate in China. Model sizes have grown from billions to hundreds of billions of parameters, while leading companies have reduced training cycles to less than 12 hours.
Traditional systems use predefined triggers—such as hard braking, sharp steering or low perception confidence—to identify short clips for review. Li Auto, for example, has disclosed more than 200 trigger conditions and can retrieve complete data from a reported vehicle issue within about one minute.
Such systems are effective for known problems but cannot easily identify unfamiliar events that were not included in the rules. Li Auto has proposed moving from a “data closed loop” to a “training closed loop,” in which engineers define the capability they want to improve and then generate or locate data specifically for that purpose.
Newer systems use vision-language models and vector databases to search driving logs by meaning rather than by fixed conditions. Engineers could submit a natural-language request for scenes involving, for example, a pedestrian rushing from behind an obstruction. The system would then identify matching moments across large video collections.
Research is also focusing on selecting smaller but more useful datasets. Studies have used risk assessments, semantic diversity and active learning to remove redundant samples. One framework reportedly achieved performance close to full-dataset training while using only 30% of the nuScenes training data.
Automakers are investing heavily in computing and data infrastructure. XPeng has disclosed a 30,000-GPU cloud cluster, a 72-billion-parameter foundation model and nearly 100 million training clips. Huawei is using generative artificial intelligence and world models to create simulated training environments, while Zeekr’s parent company, Geely, has reported cloud-computing capacity of 23.5 exaflops. Synthetic data accounted for an estimated 50% to 60% of training data in 2025, up from 20% to 30% in 2023.
The technology race is also appearing in academic publishing. Chinese automakers and suppliers published 62 autonomous-driving-related papers between August 2025 and August 2026. The research covered world models, vision-language-action systems, simulation, generative data and end-to-end planning.
BYD became the first traditional Chinese automaker to publicly release foundation-model research when it published a paper in July 2026 describing HyWorldVLA, a hybrid world-model system. The model recorded strong results on the NAVSIM benchmark, although the researchers acknowledged weaknesses in brake-light recognition and the limited viewing angle of a front-facing camera.
The paper race is partly a contest for engineering talent, technical influence and future standards. However, most of the 62 papers were produced through industry-academic collaborations, and only 24 were company-led. Publication metrics also remain difficult to compare, while turning research results into mass-produced vehicle technology can require years of engineering work.
