Chinese AI Video Models Advance as Kuaishou’s Kling 3.0 and ByteDance’s Seedance 2.0 Intensify Competition
Kuaishou Technology and ByteDance have released new generations of AI video models, Kling 3.0 and Seedance 2.0, drawing heightened market attention as competition in generative video accelerates and investors assess the emerging commercial potential of multimodal AI.
Both models show improvements over earlier versions in consistency, stability and shot composition, but the most significant upgrade is the introduction of video input capabilities. Earlier models were limited to text-to-video or image-to-video generation with basic motion controls. The latest versions allow users to upload reference videos and generate new content based on existing footage, effectively completing a multimodal input-to-video output workflow.
Market focus has shifted to how performance gains in Seedance 2.0 reshape the competitive landscape and how Kling differentiates itself. Testing across seven scenarios—including animation styles, realistic human scenes, motion performance, video-to-video conversion and lip-sync capabilities—suggests divergent positioning between the two products. Seedance is oriented toward narrative expression, while Kling is positioned for professional-grade content production, with stronger cinematic lighting, facial detail, motion control and environmental rendering.
In video-to-video tests, Kling successfully generated animated outputs from live-action input, though results were described as less dynamic than text-generated content and retained original background audio. Seedance failed to complete the same task. Pricing also reflects different strategies. A five-second 720P video costs roughly RMB 4 (about US$0.56) on Kling and RMB 2.3 (about US$0.32) on Seedance, giving ByteDance a pricing advantage, particularly for longer clips. However, Seedance currently does not support 1080P output, while Kling does.
Comparisons across global models using identical prompts indicate Kling 3.0 and Seedance 2.0 rank among the strongest performers currently available. Alibaba’s Wanxiang 2.6 was described as more cartoon-oriented with limited detail, while MiniMax’s Hailuo 2.3 produced realistic visuals but lacked synchronized audio generation. Google’s Veo 3.1 demonstrated complete baseline capabilities but produced less natural human appearances, and OpenAI’s Sora 2 was characterized as having a game-like visual style.
Pricing gaps between domestic and overseas providers remain substantial. Chinese platforms charge roughly US$0.40 for a five-second video, compared with around US$5 for Google’s model and approximately US$2.5 for Sora 2, reflecting different target user segments and cost structures.
Despite rapid technological progress, the AI video sector remains at an early commercial stage. Combined annual recurring revenue (ARR) across major video model companies totaled less than US$1 billion as of January 2026, significantly below the roughly US$20 billion ARR reported by OpenAI and US$9 billion by Anthropic in the broader AI model market. Industry ARR has grown between one and three times annually, with expansion occurring without evident revenue displacement among competitors.
Current adoption is concentrated in animation-style content, where AI tools have already replaced parts of traditional production workflows such as storyboarding, line art, motion effects and editing. Penetration into live-action short dramas, mid-length content and film production remains limited due to technical constraints, particularly resolution requirements that exceed current 720P and 1080P capabilities.
The addressable market remains substantial. China’s film box office generates roughly RMB 40 billion to RMB 60 billion annually (US$5.6 billion to US$8.4 billion), while overseas markets total about US$10 billion to US$20 billion. Including short-form video, advertising and online content production, the broader opportunity expands further, although AI penetration remains low. According to Mayor Research estimates cited in the material, China’s video production market is valued at more than US$20 billion, with the global market exceeding US$160 billion.
Industry participants argue that AI video may follow a similar trajectory to AI coding tools, where reduced production barriers expand total demand rather than merely replacing existing workflows. Growth in animation-style content consumption supports this view, with playback volumes rising sharply during the past year.
Technologically, most current video models follow a Diffusion plus Transformer (DiT) architecture, validated by multiple platforms. Alternative autoregressive approaches may emerge, potentially enabling longer video generation at higher cost. The next stage of development is expected to involve closer integration between multimodal models and so-called world models, which aim to improve physical consistency and longer-duration content generation.
For investors, the near-term outlook centers on companies combining leading models with strong distribution channels in short drama and animation ecosystems. The sector’s growth trajectory is closely tied to improvements in generation quality, commercialization progress and regulatory developments, with industry adoption still in its early phase despite rapid technological iteration.