China's AI Giants Are Rebuilding From Scratch — And Dirty Data Is to Blame

China's AI Giants Are Rebuilding From Scratch — And Dirty Data Is to Blame

A wave of pre-training restarts is exposing the industry's most expensive blind spot: years of scale-over-quality data strategies have left leading Chinese foundation model developers fighting a war they thought they had already won.

Across China's AI landscape in the past six to eight months, a quiet crisis has been unfolding. Multiple foundation model companies — including at least one that spent more than a year and tens of millions of dollars training a trillion-parameter model — have been forced to scrap their work and restart pre-training from the ground up. The culprit, in nearly every case, is not compute budgets or architectural choices. It is data. Contaminated, redundant, and poorly governed training corpora have produced models that are being outperformed on real-world benchmarks by competitors' models one-hundredth their size — a reversal that is sending shockwaves through China's AI investment community.

The reckoning is arriving at a moment when the global race for frontier AI is accelerating. For investors who have poured capital into China's model layer on the assumption that raw scale would eventually prevail, the data crisis introduces a materially new risk factor: sunk compute costs may not translate into defensible model quality, and the timeline to competitive parity with U.S. counterparts is being reset.

Dirty Data Burns Billions in Compute — and Competitive Windows

The economics of pre-training failure are brutal. Industry insiders describe a pattern that has repeated itself across multiple organizations: a model company scrapes and ingests web-scale corpora, prioritizing token volume over curation quality, and begins a training run costing tens of thousands of GPU-hours. Midway through, anomalies emerge — inexplicable repetition loops on evaluation sets, performance plateaus at 60 to 70 points on standard benchmarks while overseas rivals push past 90.

"The harm from dirty data at the foundation layer is far greater than people imagine," Wang Bin, a former senior executive at a Chinese model company and now an AI entrepreneur, told Leiphone. "It wastes every GPU, every yuan, every engineer-hour you put in — and it costs you your time window."

One case Wang described is illustrative. A mid-tier foundation model company bulk-ingested technical blog posts into its pre-training corpus. By mid-training, the team discovered that a significant portion of the content originated from content-farm mirror sites — machine-generated pages in which individual paragraphs were duplicated tens of thousands of times. The model began exhibiting pathological repetition behavior on evaluation sets. By the time the team traced the root cause, tens of thousands of GPU-hours had been consumed. The entire training run was abandoned.

Zhao Xiang, a data specialist at a top-tier foundation model company, frames the early-stage mentality that produced these outcomes: "The prevailing view was that more data was fine, even if quality was slightly lower — the model has its own conversion capacity and robustness." That assumption, he told Leiphone, proved catastrophically wrong at scale.

Misaligned KPIs Turned Data Teams Into Quantity Machines

The problem was not purely technical. Organizational incentive structures amplified the damage. Wang Yunhe, founder of Jiyuan Lüdong, has publicly argued that data teams must hold meaningful authority over pre-training teams — and vice versa — to prevent the siloed dysfunction that volume-only metrics create.

That dysfunction materialized visibly at Tencent's Hunyuan team, according to media accounts. Annotation rules were insufficiently specified, yet acceptance thresholds were set high, creating a perverse incentive: annotators produced large volumes of technically compliant but practically unusable labeled data. Benchmark contamination and redundant data further degraded corpus integrity. The team ultimately dismantled its data operation and rebuilt it from scratch.

"If you evaluate a data team purely on volume, you get absurdity," Zhao Xiang said. "Foundation models compete on information density and cleanliness, not absolute corpus size. Hundreds of terabytes of dirty data fed into training will make your model dumber, or crash the run entirely."

Internet-Era Data Moats Dissolve Under Foundation Model Demands

One of the more counterintuitive findings from this data crisis is the effective obsolescence of the data assets that China's internet giants spent two decades accumulating. Platforms commanding dominant positions in local services, social networking, gaming, e-commerce, and search — assets once considered impenetrable competitive moats — are proving largely irrelevant to foundation model pre-training.

"Google has more data across more verticals than anyone, yet that clearly hasn't translated into model supremacy," said Xu Dong, an algorithm lead at a foundation model company. "Anthropic and OpenAI started with zero proprietary data accumulation and are still pushing harder. The same dynamic holds domestically."

The reason is structural. Foundation models optimize for general capability, and vertical domain data — however voluminous — contributes marginally to general benchmarks. The companies that have moved fastest are those that built rapid, high-quality data pipelines, not those with the largest legacy corpora.

Budget Allocations Reveal Data's Chronic Underfunding

Despite the industry's rhetorical emphasis on data quality, actual budget allocations tell a different story. Zhou Jun, a business director at an AI data supplier, shared an industry rule-of-thumb breakdown with Leiphone: of every RMB 100 (US$13.89) invested in AI training, approximately RMB 40 goes to compute, RMB 30 to talent, RMB 20 to marketing, and just RMB 10 to data.

The gap becomes even more striking when examining expert data procurement. For high-difficulty reasoning tasks of the type used in benchmarks like HLE (Humanity's Last Exam), U.S. buyers typically pay US$10,000 to US$20,000 per annotated example. Chinese model companies, Zhou said, are offering RMB 1,000 to RMB 2,000 (US$139 to US$278) for comparable work — a price differential that directly constrains the quality and quantity of expert-labeled data available to domestic developers.


API Relay Stations Become Covert Data Harvesting Operations

Supply constraints have driven some players toward less transparent acquisition strategies. Zhou Jun described a practice that has become an open secret in the industry: certain model companies operate API relay or token-forwarding services — ostensibly commercial middleware products — that intercept and retain the interaction traces of users calling third-party frontier models.

When enterprise or developer customers route API calls through these intermediaries to access models from providers such as Anthropic or OpenAI, their queries and the corresponding model outputs pass through the relay layer and are logged. The resulting dataset — real task requests paired with high-quality model responses — is precisely the distribution that model companies need for post-training alignment and distillation, and it is far more valuable than scraped web content.

Model routing startups are also eyeing this data asset. Major model companies operate their own routing platforms, but because these platforms preferentially direct traffic to proprietary models, they capture limited interaction data from competing frontier systems — and their user bases remain constrained as a result.

"Over the past few years, companies have taken every conceivable approach to acquiring AI training data," Zhou said. "How long any one of these channels can sustain a competitive advantage is genuinely unclear."

The Data Flywheel Thesis Hits Three Structural Blockages

The strategic logic that has animated significant investment in China's AI sector — that stronger models attract higher-quality user interactions, which feed back into training to produce still-stronger models — remains theoretically sound. In practice, almost no Chinese model company has successfully closed this loop at scale.

Three structural blockages are impeding flywheel formation. First, ongoing organizational restructuring at major technology groups means that data walls between product divisions and model training teams remain largely intact, with cross-functional data sharing still in early stages. Second, the talent market for data engineering roles — data cleaning specialists, pipeline architects, data researchers — has only recently begun to price these skills appropriately, having historically classified them as support functions. Multiple AI recruiters confirmed to Leiphone that these roles are among the most difficult to fill in 2026.

Third, and most consequentially, Chinese model companies face a structural deficit in the long-horizon trajectory data that is increasingly critical for post-training and agent capability development. The two domestic scenarios most likely to generate this data — AI coding and AI-assisted office productivity — are both constrained. In coding, offshore models including Anthropic's Claude and OpenAI's Codex hold substantial market share among Chinese developers, meaning the interaction traces flow to foreign training pipelines. In office productivity, domestic products remain in early stages, limiting the complexity and volume of task trajectories they can capture.

A Long War With No Shortcuts — and More Restarts Ahead

The pre-training restarts of the past six to eight months represent the first wave of an industry-wide reckoning, not its conclusion. Multiple senior practitioners told Leiphone that additional "rebuild from scratch" events are probable within the next 12 to 24 months as companies that deferred data governance work encounter the consequences at later training stages.

The strategic implication for investors is significant. Compute capacity is purchasable. Talent is recruitable, at a price. High-quality training data accumulates on a different timeline — one governed by organizational discipline, pipeline maturity, and the slow accretion of genuine user interaction at scale. Companies that have not yet built the internal infrastructure to convert product usage into training signal are not simply behind on a technical metric; they are operating without the raw material that the next phase of model development requires.

The companies most likely to emerge from this data correction cycle with durable advantages are those that can first repair the organizational plumbing — breaking down internal data silos, properly incentivizing data quality over volume, and securing the long-horizon expert and trajectory data that pre-training-era corpora cannot supply. That work is unglamorous, expensive, and slow. By most accounts, for the majority of China's model companies, it has barely begun.

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe