Alibaba's Qwen3.7-Plus Tops GPT-5.4 in Screen Tasks, Builds Apps Autonomously

Alibaba's Qwen3.7-Plus Tops GPT-5.4 in Screen Tasks, Builds Apps Autonomously

Alibaba has deployed a multimodal artificial intelligence agent capable of autonomously navigating graphical interfaces and engineering complete software applications, challenging OpenAI’s GPT-5.4 in the escalating 2026 race for enterprise automation.

Released by the company’s Tongyi Qianwen cloud division, Qwen3.7-Plus marks a critical transition in AI utility—shifting from passive text generation to active workflow execution. In a demonstrated 11-hour continuous loop, the hybrid-agent system independently coded, tested, and deployed functional applications, signaling a structural shift in software development economics.

Initial developer feedback has centered on the model’s commercial viability for "digital worker" deployment. Industry analysts note that Qwen3.7-Plus effectively bridges the gap between visual perception and action, integrating Graphical User Interface (GUI) navigation with Command Line Interface (CLI) execution to operate existing enterprise software exactly as a human employee would.

Bridging Vision and Execution Disrupts Software Workflows

The strategic value of Qwen3.7-Plus lies in its end-to-end autonomous execution capabilities. During a sustained 11-hour session, the hybrid-agent system autonomously managed the complete research and development lifecycle of an English vocabulary application. The AI generated over 10,000 lines of code, executed more than 1,000 agent calls, and handled automated deployment, GUI testing, and version iteration without human intervention.

Further demonstrating its capacity to handle complex data pipelines, the model successfully replicated a native macOS Stocks application. It processed UI layouts, generated SwiftUI source code, and integrated the LongBridge real-time market API to fetch live financial data. The agent subsequently executed and passed 10 functional validation tests, including multi-period chart switching and real-time data loading, proving its ability to synthesize visual design with backend data architecture.

Topping US Incumbents Validates Commercial Viability

Benchmark data places Alibaba’s latest iteration ahead of major US competitors in specific multimodal and operational tasks. On the ScreenSpot Pro benchmark—a critical metric for measuring an AI's ability to understand and interact with screens—Qwen3.7-Plus scored 79.0, eclipsing OpenAI’s GPT-5.4 (67.4) and Google’s Gemini 3.1 Pro (68.1). In the AndroidWorld mobile manipulation evaluation, it achieved an 81.0, outperforming both Gemini 3.1 Pro (70.7) and Anthropic’s Opus-4.6 Max (62.0).

The model also demonstrated exceptional proficiency in translating visual inputs into code, scoring a leading 85.9 in Chart Recognition (CharXiv). It can convert static images, UI screenshots, and UI/UX reference designs directly into scalable vector graphics (SVG) or functional web prototypes.

However, the architecture still faces friction in highly complex reasoning environments. Qwen3.7-Plus scored 77.7 on the SWE-Verified software engineering benchmark, trailing Opus-4.6 Max (80.8), and underperformed GPT-5.4 in the Hard Logic Evaluation (HLE) metric.

Intensifying Domestic Rivalry Accelerates Monetization

The release timing underscores a hyper-competitive domestic AI landscape in 2026. Alibaba’s launch occurred just one day after Chinese AI unicorn MiniMax debuted its M3 model. While MiniMax focuses on open-source deployment and a massive 1-million token context window optimized for long-thread autonomous academic research, Alibaba is pursuing a closed-weight, API-driven monetization strategy.

Priced at US$0.4 per million input tokens and US$1.6 per million output tokens, Qwen3.7-Plus is positioned squarely at enterprise clients looking to automate routine software engineering and GUI-based administrative tasks. By consolidating "seeing, thinking, coding, and acting" into a unified foundation model, Alibaba is betting that the immediate future of AI monetization lies not in raw parameter scale, but in seamless integration with existing desktop and mobile operating systems.

Related Coverage:

Alibaba's Qwen3.7-Max Tops China AI Rankings, Closes Gap With Claude and GPT in Agentic Benchmarks

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe