Qwen Office Beats Claude and Codex as Harness Emerges as AI's New Moat

Qwen Office Beats Claude and Codex as Harness Emerges as AI's New Moat

Qwen Office scores 95/100 in Wall Street's first systematic head-to-head test of eight Chinese and U.S. AI agents, beating rivals from Anthropic and OpenAI while running on a model ranked fourth in raw intelligence — a result that forces a fundamental reassessment of where enterprise AI value is actually created.

Jefferies, the New York-based investment bank, published findings this week from what it describes as the most rigorous cross-border AI agent benchmark conducted to date, testing eight commercial products across five enterprise workflow tasks. The headline finding: the quality of an agent's "harness" — the orchestration layer governing instructions, context, tooling, guardrails, feedback loops, and governance — is a more reliable predictor of real-world performance than the intelligence score of the underlying model. The report lands as enterprise software buyers face a crowded, often opaque market, and as investors attempt to separate durable competitive moats from distribution-driven growth stories.

Market reaction among AI-adjacent equities was immediate. The report's framing challenges the prevailing assumption that model capability is the primary determinant of enterprise agent value, a thesis that has underpinned premium valuations for foundation model providers throughout 2025 and into 2026.


Harness Engineering, Not Model IQ, Drives Enterprise Agent Outcomes

Jefferies decomposed the agent stack into two discrete components: the model (responsible for reasoning) and the harness (responsible for management). The harness was further segmented into six functional layers — instructions, context, tools, boundaries, feedback, and governance — each addressing coordination problems the model itself cannot resolve autonomously.

The bank's controlled-variable evidence is striking. Holding the model constant and varying only the harness, Claude Opus 4.6 scored between 58.0% and 76.4% on Terminal-Bench 2.0 — an 18.4 percentage point spread attributable entirely to harness engineering. Gemini 3 Pro showed a 13.4-point spread under identical conditions. The implication for enterprise buyers is direct: procurement decisions anchored solely on model benchmarks are systematically mispriced.

Jefferies applied a weighted decomposition — 60% model, 40% harness — to back-calculate an "implied harness score" for each product. Alibaba's Qwen Office posted the highest implied harness score in the entire field, surpassing both Anthropic's Claude Cowork and OpenAI's Codex.


Qwen Office Executes a Quiet Reversal at the Top of the Leaderboard

Final scores across the eight agents tested: Qwen Office 95, Claude Cowork 94, Codex 92, Kimi Work 86, Doubao 77, MiniMax Code 71, with Tencent's Workbuddy and Google's Gemini Spark tied at 66.

The Qwen Office result is analytically significant precisely because it is counterintuitive. Its underlying model, Qwen 3.8 Max, carries a model intelligence score of 56 — fourth among the eight products tested. Claude Opus 5, backing Claude Cowork, scores 61; GPT-5.6 Sol, backing Codex, scores 59. Qwen Office's harness engineering closed a five-to-six point model intelligence gap and converted it into a one-point overall lead. That is a material swing.

Pricing amplifies the strategic advantage. Qwen 3.8 Max is available via API at approximately $1.10 per million tokens, versus $4.40 for GPT-5.6 Sol and $3.90 for Opus 5 — a cost differential of roughly 70% to 80%, consistent with the broader Chinese model pricing dynamic driven by mandatory efficiency optimization under export controls and intense domestic competition. For agent workflows that consume large volumes of inference tokens, this pricing gap directly expands the addressable enterprise market.

Task-level analysis reveals a structural divergence between Chinese and U.S. agents. Chinese products, including Qwen Office and Doubao, outperformed on the marketing poster generation task (Task 5), where Claude Cowork and Gemini Spark both failed. U.S. agents held a clear advantage on browser control tasks (Task 2), where Workbuddy and MiniMax Code underperformed. The divergence likely reflects optimization toward domestic use cases and software environments on the Chinese side.


Workbuddy's Anomaly Exposes the Gap Between Distribution and Engineering

The most analytically complex finding in the Jefferies report concerns Tencent's Workbuddy. By traffic metrics, Workbuddy leads the Chinese market: monthly visits reached approximately 21 million in June 2026, the highest among domestic peers. By harness engineering score, it ranks last among the five Chinese agents tested.

Jefferies attributes the divergence to three factors independent of harness quality. First, deep integration with Tencent's proprietary ecosystem — Tencent Docs, IMA, and WeCom — provides distribution advantages unavailable to standalone products. Second, Workbuddy operates on a model-agnostic architecture, allowing users to swap between third-party models including Kimi K3, DeepSeek V4, and GLM 5.2, effectively borrowing model strength rather than building it. Third, Tencent's marketing investment has been substantial.

The Jefferies verdict is nuanced: Workbuddy wins on distribution in the near term, but its 21 million monthly active users represent a compounding data asset. Every agent interaction generates a trace — tool calls, failure modes, human corrections — that feeds reinforcement learning pipelines and harness improvement cycles. The question for investors is whether Tencent converts that data flywheel into harness engineering parity before competitors with stronger harness foundations capture enterprise accounts.

ByteDance's Trae and Workbuddy both pursue model-agnostic harness strategies, a deliberate architectural choice that Jefferies frames as a structural advantage in a market where model commoditization is accelerating. The counterargument — that model-agnostic harnesses are more vulnerable to being displaced by vertically integrated competitors — is not addressed in the report.


China's Harness Landscape: Four Structural Leads, Four Structural Constraints

Jefferies mapped the competitive dynamics between Chinese and U.S. harness approaches across eight dimensions.

On the advantage side: super-app integration (Chinese agents embedded directly in DingTalk, Feishu, and WeCom collapse the tool-connector problem that U.S. agents solve via third-party connectors); model-agnostic flexibility (enabling cost and capability optimization across a competitive open-source model market); token pricing 70% to 80% below U.S. equivalents (making compute-intensive agentic workflows economically viable at scale); and iteration velocity (large domestic user bases generate failure-mode data faster, compressing harness improvement cycles).

On the constraint side: the underlying model gap persists, and weaker models require proportionally better harness engineering to compensate — a structural tax on Chinese agent builders. Enterprise software monetization remains structurally difficult: SME price sensitivity and large state-owned enterprise preference for custom project contracts over subscription models suppresses recurring revenue and, by extension, the capital available for sustained harness R&D. International expansion is constrained by optimization for domestic ecosystems and user behaviors incompatible with globally standardized software stacks. And high-end compute scarcity — a direct consequence of U.S. export controls — creates service reliability risks and caps inference capacity precisely as agentic workflows demand orders of magnitude more compute than conversational AI.


Harness Compounds Into Switching Costs, Data Moats, and Willingness to Pay

The Jefferies framework concludes with an investment-relevant observation: enterprise customers are not purchasing intelligence. They are purchasing a complete, integrated product in which the harness — carrying accumulated workflow history, memory, connectors, skills, and automations — is the durable asset. The model is interchangeable; the harness is not.

This reframes the competitive moat question. As model capabilities converge — a trend already visible in the narrow score spreads at the top of the Jefferies benchmark — harness quality becomes the primary differentiator. Switching costs accrue to the harness layer, not the model layer. The data flywheel (more users → more traces → better reinforcement learning → better harness → more users) favors incumbents with large, engaged enterprise user bases over new entrants with superior models.

For enterprise AI buyers evaluating procurement in the second half of 2026, the Jefferies findings suggest a practical reorientation: benchmark harness engineering — instruction clarity, context management, tool reliability, boundary enforcement, feedback loop quality, and governance controls — as rigorously as model intelligence scores. The two are not correlated, and the former is more predictive of actual workflow outcomes.

Related Coverage:

Alibaba Releases Weights for 2.4T-Parameter Qwen3.8, Escalating Open-Source AI Arms Race

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe