Alibaba touts “HappyHorse” video model after anonymous benchmark win, signaling broader multimodal AI push

Alibaba touts “HappyHorse” video model after anonymous benchmark win, signaling broader multimodal AI push

Alibaba Group Holding has moved to formalize an unexpected win in generative video, acknowledging it built “HappyHorse 1.0” after the model appeared anonymously on Artificial Analysis’ Video Arena and reached No.1 in both text-to-video and image-to-video within 48 hours—an outcome that reshapes investor expectations for the company’s AI execution beyond chatbots.

The company said HappyHorse 1.0 was developed inside its ATH unit and is currently in closed beta, with an API expected to open “soon.” The disclosure reframes what had looked like a third-party breakthrough into a deliberate move by Alibaba to test competitiveness in public benchmarks before attaching its brand.

The result also disrupts a video-model pecking order that had largely been treated as stable in China, with ByteDance’s Seedance widely viewed as the frontrunner and Kuaishou Technology’s Kling seen as the closest challenger. Alibaba’s new placement suggests the gap is narrowing—and that Chinese incumbents are now willing to use global, model-to-model arenas as a marketing and recruiting tool.

Revealing Internal Parallel Teams Signals Faster Iteration

HappyHorse’s emergence adds a second visible line of effort to Alibaba’s video strategy, which had been publicly associated mainly with Tongyi Wanxiang. On April 7, Alibaba released Wan2.7-Video, highlighting capabilities across complex motion, audio-video synchronization, longer-form generation and video editing—features that typically require heavier training data pipelines and more compute-efficient inference.

Alibaba’s latest internal structure puts these efforts under the same ATH group but on different tracks: Tongyi Wanxiang sits within the Tongyi foundation-model unit, while HappyHorse comes from an AI innovation unit that is positioned closer to product scenarios. For investors, that separation reads as an “internal competition” model aimed at compressing development cycles—while reducing key-person risk in top-tier multimodal training.

Alibaba also said it plans to launch another multimodal model distinct from HappyHorse in the near term, indicating it is treating video and multimodal systems as a portfolio rather than a single flagship bet.

Stacking External Investments Builds a Second Line of Defense

Alibaba is pairing in-house development with external capital to secure optionality in model architectures and go-to-market routes. Around the same period as HappyHorse’s benchmark rise, Alibaba led Shengshu Technology’s Series B round of RMB 2 billion (US$278 million), the company behind the Vidu multimodal model, which has ranked in Artificial Analysis’ video leaderboard top 10.

Alibaba has also been cited as a leading investor in Aishi Technology, another Chinese AI video generation player. The combined pattern—multiple internal teams plus stakes in fast-moving startups—mirrors a “two-layer insurance” strategy: claim technical leadership with proprietary models while reserving ecosystem influence through minority ownership.

For the supply chain, that approach can translate into steadier demand signals for GPU capacity, data labeling, model evaluation tooling and cloud inference optimization—because Alibaba is effectively funding more than one roadmap that consumes compute.

Repositioning Video as a Compute-Heavy Gateway to Agents

The strategic logic goes beyond making clips. Video models stress test temporal consistency, physical motion, camera control, audio-visual alignment and inference efficiency—dimensions that are harder to fake with parameter scaling alone. As Alibaba shifts AI priorities toward “agents” that connect into its commerce and enterprise software footprint, video becomes a gateway to video understanding, multimodal agents and new interaction patterns.

Shengshu’s financing plan also points to “world model” work, reinforcing that Alibaba is linking video generation to longer-horizon multimodal reasoning rather than treating it as a standalone creative app category.

That linkage matters because Alibaba has already committed to heavy infrastructure spending. Early 2026, the company announced it would invest at least RMB 3.8 trillion (US$528 billion) over three years in AI and cloud infrastructure. Video generation and multimodal systems are among the few application layers capable of continuously absorbing that level of compute—turning model leadership into cloud utilization and, eventually, pricing power in enterprise AI services.

Forcing Rivals to Respond Raises the Stakes in China’s Video Race

By claiming an anonymous benchmark winner and signaling multiple releases, Alibaba is pressuring ByteDance and Kuaishou to defend their perceived lead not only with demos but with third-party rankings, APIs and enterprise-grade reliability. If HappyHorse’s API opens quickly, it could accelerate competition in downstream integrations—advertising creative, e-commerce content, short-video tooling and customer service—markets where Alibaba has distribution advantages through Taobao and Tmall, Alibaba Cloud and its merchant ecosystem.

The near-term question for markets is execution: whether Alibaba can convert leaderboard visibility into stable developer adoption and cloud revenue, while managing the higher inference costs and safety constraints that come with video.

Related Coverage:

Alibaba Unveils Qwen3.5-Omni With 215 SOTA Wins, Disrupting Global Multimodal AI Race

Alibaba Cloud Opens Wan2.7-Image API, Targets Monetizable AI Images

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe