China's AIGC Entertainment Race: 12 Companies, One Market

China's AIGC Entertainment Race: 12 Companies, One Market

The question in China's AI-generated content sector has shifted. It is no longer "which model produces the most impressive video clip." It is now "who can build a reliable, controllable, and monetizable content production system at scale."

That distinction matters enormously—and it explains why the competitive landscape looks so different in 2026 than it did just two years ago.


What Is AIGC Entertainment, and Why Does It Matter Now?

AIGC—AI-Generated Content—refers to media produced wholly or substantially by artificial intelligence models, covering video, audio, images, music, and interactive assets. In the entertainment context, this means AI-generated short dramas, animated comics, game assets, advertising creatives, and eventually feature-length productions.

The reason this matters now is structural, not cyclical. Three forces are converging simultaneously:

Model capability has crossed a practical threshold. Video generation has progressed from producing a few seconds of passable footage to supporting multi-character performance, complex camera movement, synchronized audio, and multi-modal inputs (text, image, audio, and reference video combined). What was a novelty in 2023 is becoming a production tool in 2026.

The cost curve for content production is bending sharply. Short dramas, animated comics, and advertising materials—high-volume, short-lifecycle formats—are the first categories where AI can demonstrably compress production time and cost. One workflow system in China has publicly targeted a reduction in per-episode production time to between 30 minutes and one hour.

Commercial revenue has appeared, not just user metrics. Kuaishou's Kling AI reported quarterly revenue exceeding RMB 650 million in Q1 2026, a year-on-year increase of more than 300%, with an annualized revenue run rate approaching USD 500 million. MiniMax reported USD 79 million in revenue for 2025, growing 158.9% year-on-year, with over 70% coming from international markets. These are no longer funding stories—they are businesses with measurable cash flows.


How the Production Chain Actually Works

Understanding the competitive dynamics requires mapping the full value chain, because different players are positioned at different points along it.

Upstream: Foundation models. This layer covers image generation, video generation, audio synthesis, music generation, and 3D asset creation. Raw model capability—resolution, motion coherence, character consistency, audio-visual synchronization—is determined here.

Midstream: Production systems. A single impressive shot does not make a drama series. The midstream layer covers scriptwriting, character asset management, storyboarding, shot sequencing, scene generation, audio integration, and editing. The critical challenge at this layer is consistency: can the same character look the same across 200 shots? Can a production team modify one scene without regenerating everything else? Workflow orchestration, not model benchmarks, is the defining capability here.

Downstream: Distribution, IP, and monetization. Who owns the story IP? Who controls the distribution platform? Who has the paying users? Short drama platforms, short video feeds, gaming ecosystems, advertising networks, and international streaming services all represent potential endpoints—but access to them is not equally distributed.

The companies competing in this space have staked out very different positions across these three layers.


The Major Players and Their Strategic Logic

Internet Giants: Competing with Ecosystems

ByteDance is currently the closest to a fully integrated vertical. Its model stack includes Seedream (image), Seedance (video), and multi-modal audio and voice capabilities. Its product layer includes Doubao, Jianying (CapCut), and Jimeng AI. Its content and distribution layer includes Douyin, TikTok, Fanqie Novel, and Honguo short drama platform.

Seedance 2.0, released in February 2026, introduced a unified audio-video joint generation architecture supporting text, image, audio, and video inputs, with an emphasis on complex motion, camera direction, and native audio-visual synchronization. The subsequent Seedance 2.5 extended these capabilities further.

ByteDance's structural advantage is not any single model ranking—it is the closed loop: IP from Fanqie, visual production via Jimeng and Jianying, distribution via Douyin and Honguo, and user behavior data feeding back into content selection. For most companies, AI video is a new product line. For ByteDance, it may become the infrastructure for rebuilding its entire content supply system.

Kuaishou has taken a more concentrated approach, positioning Kling AI as a standalone global creative production platform. Kling has moved beyond consumer viral content into professional production—it was involved in virtual scene and visual effects work for the television drama Taiping Nian. Kuaishou is packaging model capability directly into subscriptions, credit systems, API access, and enterprise services. Its existing short video ecosystem and creator network provide a natural seed user base.

Alibaba is building around Tongyi Wanxiang (Wan), its image and video generation system. Wan 2.6 supports reference-video generation, multi-person dialogue, multi-shot narrative, shot control, and native audio, with single-generation lengths up to 15 seconds. The model Wan2.7-Video targets creative freedom specifically. Alibaba's real leverage, however, lies in Alibaba Cloud, Youku, Taobao advertising, and its merchant ecosystem—Wan can serve both professional film production and high-volume commercial video generation. Alibaba's investment in PixVerse developer AISphere signals that it is supplementing internal model development with external talent acquisition.

Tencent is focused on gaming and IP. Its Hunyuan model family covers image, video, 3D, and world modeling. Hunyuan 3D can generate editable 3D assets from text, images, or multi-view inputs, and is extending toward spatial content with physics simulation and character navigation. Tencent Games has launched a Hunyuan game visual generation platform, the GiiNEX game AI engine, and a UGC game creation platform codenamed "Craft." The competitive logic is distinct: games require models that understand three-dimensional space, character behavior, and player feedback—not just linear narrative. Tencent is competing for the "generatable, interactive digital world."

Baidu is entering through enterprise video generation and search marketing. Its MuseSteamer system covers environment audio, character voice, multi-person dialogue, and long-form video production, accessible to enterprise clients via the Qianfan platform. Without a content community comparable to Douyin or a gaming empire comparable to Tencent's, Baidu's most likely near-term monetization is in marketing video, knowledge content, digital human broadcasting, and enterprise creative assets.


Specialist Platforms and Model Startups: Competing for Vertical Depth

360 has built what it calls the Nami comic drama production pipeline—an end-to-end workflow that integrates script decomposition, character asset management, AI storyboarding, image generation (including external models such as Seedance 2.0), and post-production. The stated target is a per-episode production time of 30 minutes to one hour, with a shot generation success rate above 90%.

This represents a distinct strategic logic: 360 does not need to lead on any single model benchmark. It needs to orchestrate multiple models, standardize workflows, and maintain character consistency across episodes. The model determines the ceiling of any individual shot; the pipeline determines how many episodes a studio can produce per day.

SenseTime's Seko is evolving in a similar direction, covering story creation, storyboarding, character design, shot organization, and final delivery. By July 2026, the platform reported over one million creator users and 1,500 enterprise clients. SenseTime is combining its visual AI heritage with agent-based orchestration to position Seko as an "AI video dream factory" rather than a generation tool.

Zhipu AI entered video generation relatively early with its Qingying model. Qingying 2.0 supports 10-second, 4K, 60fps generation with audio effects, accessible directly within the Zhipu Qingyan product. Zhipu's differentiation is the combination of general large language model capability with video generation—enabling creative ideation, script writing, and prompt engineering within a single system. However, in the entertainment vertical specifically, Zhipu currently operates more as an infrastructure provider than a platform with a closed commercial loop.

Stepfun emphasizes multi-modal foundations. Its Step-Video-T2V, a 30-billion-parameter model, has been open-sourced under a permissive MIT license. Step-Audio covers speech, emotion, dialect, singing, and character roleplay. The open-source strategy accelerates developer ecosystem growth but raises a structural question: when a model can be freely deployed by any enterprise, where does the value ultimately accrue—to the model company, the cloud provider, or the application platform with user relationships?

MiniMax has answered that question by building AI-native consumer products directly. Hailuo AI handles video, MiniMax Audio covers voice synthesis, a dedicated Music model covers music generation, and Talkie-AI explores AI character companionship. With 2025 revenue of USD 79 million (up 158.9%), over 70% from international markets, and a cumulative user base exceeding 236 million individuals plus 214,000 enterprise clients and developers, MiniMax has demonstrated that a Chinese company can charge international users directly for AI entertainment products. The risk is proportional to the success: global copyright exposure—covering training data, celebrity likeness generation, and user-created derivative content—grows with international scale.

Kunlun Tech is positioning itself as an AI entertainment group. SkyReels targets AI short drama production, Mureka covers AI music, and DramaWave enters overseas short drama consumption directly. SkyReels-V1 was disclosed in annual filings as trained on film and television data, with a character parsing system to strengthen performance and shot generation. The strategic ambition is to simultaneously own the production tool and the content platform—but model development, content production, and international distribution are all capital-intensive businesses, and whether the three lines reinforce or cannibalize each other will determine the strategy's viability.

AISphere represents the vertically focused AI video startup model. PixVerse serves international markets; Paime AI serves domestic users. By September 2025, PixVerse had reached 175 countries with over 100 million users, and the company disclosed annual recurring revenue exceeding USD 40 million. A large Series C round in 2026 is being directed toward video foundation models, real-time world models, and international growth. AISphere's competitive advantage is product iteration speed and viral format design—effects templates, character transformation, and low-barrier generation spread rapidly on social platforms. The risk is that viral formats are easily replicated, which is why AISphere is extending toward real-time interactive video and world models.


Three Competitive Layers—and Why They Matter Differently

Mapping these twelve companies reveals that the competition is actually occurring on three distinct levels, with different dynamics at each.

Layer 1: Model capability. Video generation has advanced from text-to-video and image-to-video, to integrated audio-visual generation, multi-character performance, complex camera work, and multi-modal reference inputs. ByteDance, Kuaishou, Alibaba, Baidu, MiniMax, and AISphere are all competing intensely here. But raw model performance is increasingly difficult to sustain as a long-term moat. Model update cycles are compressing; a benchmark lead may last only weeks.

Layer 2: Production systems. Short dramas, animated comics, games, and advertising campaigns are not single shots—they are sequences of dozens or hundreds of shots requiring character consistency, narrative continuity, reusable asset libraries, and selective editing capability. 360's Nami pipeline, SenseTime's Seko, and Tencent's game AI platforms all represent the shift from model competition to workflow competition. This layer has higher switching costs and is harder to replicate quickly.

Layer 3: IP, users, and distribution. This is where the final commercial competition is decided. ByteDance controls novel IP, short drama distribution, short video feeds, and editing tools. Kuaishou has a creator ecosystem and demonstrated revenue scale. Tencent owns gaming IP, animation IP, and long-form entertainment IP. Alibaba controls e-commerce advertising, Youku, and cloud infrastructure. Kunlun Tech, MiniMax, and AISphere are building international user bases. By contrast, Zhipu AI, Stepfun, and Baidu are more dependent on API revenue, enterprise contracts, and industry partnerships to convert model capability into stable order flow.


What AI Will—and Will Not—Disrupt in Entertainment

The first wave of AI disruption in entertainment is unlikely to be a fully AI-produced feature film. It will be the systematic displacement of mid-to-low-cost content formats: novel adaptation posts, animated comics, advertising creatives, game assets, short video templates, and overseas short dramas. These formats share common characteristics: high volume, short content lifecycles, high tolerance for imperfection, and calculable return on investment.

As production costs fall across the industry, content supply will increase substantially—and oversupply will intensify. When that happens, the scarce resources shift. The question of what story is worth tellingwhich character will be remembered, and what content users will actually pay for becomes more important, not less.

AIGC will not transform entertainment into a pure technology industry. When every company can generate high-quality visuals, the technical differentiation converges. IP ownership, aesthetic judgment, organizational capability, distribution efficiency, and copyright governance become the differentiating factors again.

The competition today appears to be between models. The final outcome will be determined by the content industry.


Key Variables to Watch

Character consistency at scale. The ability to maintain a coherent character across a full drama series—not just a single scene—remains technically unsolved at production quality. Whichever workflow system achieves reliable consistency first gains a structural advantage in professional content production.

Copyright and training data liability. As AI-generated content becomes commercially significant, legal exposure around training data sourcing, celebrity likeness generation, and user-created derivative works will increase. This is particularly acute for companies with international revenue, where legal frameworks are more developed and enforcement is more active.

Open-source versus closed-model economics. Stepfun's decision to open-source Step-Video-T2V illustrates the tension: open models build developer ecosystems but commoditize the model layer, potentially shifting value to cloud providers and application platforms. How this plays out will affect the long-term economics of model-focused startups.

International monetization. MiniMax, AISphere, Kunlun Tech, and Kuaishou are all generating meaningful international revenue. Whether Chinese AI entertainment products can sustain global user growth while managing copyright, content moderation, and geopolitical risk is the defining question for the sector's international ambitions.

Consolidation timeline. The current field of twelve significant players is almost certainly too large to persist. As the workflow layer matures and distribution advantages compound, the industry will likely consolidate around a smaller number of companies that have successfully integrated model capability, production systems, and content distribution into a coherent commercial loop.

Related Coverage:

ByteDance's AI Pivot: Why China's Tech Giant Is Betting Its Future on Enterprise Productivity

Kling AI's Rise: How Kuaishou Built China's First Commercially Viable Video Generation Model

Alibaba Cloud's Domestic AI Supernode Goes Live, Igniting China's Compute Infrastructure Race

China's AI Cloud Wars: How Tencent, Alibaba, and Baidu Are Adapting Their Strategies

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe