Zhipu AI open-sources GLM-5.1, raises prices 10% as China’s models shift from price war to performance premium
Zhipu AI is betting that benchmark-leading code performance can buy pricing power: the company on April 8 open-sourced its flagship GLM-5.1 model and lifted prices by 10%, positioning it as a rare open model that can match—or beat—top closed rivals in real-world software engineering tests.
The release marks a sharp pivot for China’s large-model market, which spent much of the past year cutting prices “by more than 90%” to gain share, according to the provided material. By moving the other way—raising prices while emphasizing long-horizon “project completion” rather than chat—Zhipu is trying to re-anchor enterprise buying decisions around delivered engineering output, not token discounts.
Early signals of that repositioning show up in third-party pricing comparisons. The model-aggregation platform OpenRouter shows GLM pricing increased again by 10%, and after the adjustment GLM-5.1’s cached token price for coding is close to Anthropic’s Claude Sonnet 4.6, according to the material.
Beating benchmarks resets open-source competitiveness
GLM-5.1’s central claim is that it breaks the long-standing capability gap between open models and elite closed systems. In SWE-bench Pro—a benchmark built on industrial-grade software tasks drawn from real GitHub repositories—GLM-5.1 set a new global score and recorded the first reported overtake of Claude Opus 4.6 in that test, according to the material.
Across three code-centric benchmarks—SWE-Bench Pro, Terminal-Bench 2.0 and NL2Repo—GLM-5.1 ranked third globally on the combined average, while placing first among China models and first among open-source models, the material said. For investors and buyers, that framing matters because code generation and tool use are increasingly treated as the clearest proxy for whether a model can move from “answering” to “shipping.”
Extending task duration pushes AI toward “project delivery”
Zhipu is also using a different yardstick: how long a model can work autonomously on a single assignment. The company said GLM-5.1 can sustain work for up to eight hours per task, with planning, execution and self-improvement loops, putting it among a small set of models with that “8-hour” capability and one of the few globally outside Claude Opus 4.6, based on the material.
The emphasis aligns with “Task-Completion Time Horizon,” a metric proposed by AI safety research organization METR in March 2025, which focuses on the duration of tasks a model can complete independently. The material said METR’s research suggests the time horizon doubles every seven months; MIT Technology Review called the curve “the most important chart in AI,” and Sequoia Capital cited it in early 2026 when declaring “this is AGI.”
Zhipu attributed the longer-horizon behavior to training changes such as expanding the training window for task processes and optimizing tool use, enabling an “experiment→analysis→optimization” closed loop, according to the material.
Raising prices signals a shift from token subsidy to value-based monetization
The 10% price increase is the clearest commercial signal in the release. With coding prices nearing Anthropic’s level on OpenRouter, Zhipu is effectively arguing it can deliver comparable engineering value and thus charge accordingly, the material said.
Zhipu CEO Zhang Peng said at Zhongguancun Forum that long-term reliance on low-price competition is unhealthy for the industry, and that the adjustment aims to return prices to a “normal commercial value range.” He added that long-horizon tasks can consume 10 to 100 times the tokens of a simple answer, making the change a “natural result” of value shifting, according to the material.
Demonstrating eight-hour autonomous work reframes enterprise ROI
To illustrate the “project completion” pitch, Zhipu described an “8-hour build a Linux desktop from scratch” run: GLM-5.1 executed more than 1,700 steps, produced a meaningful result within 20 minutes, and delivered a functional desktop system with a window manager, status bar, apps, VPN manager, Chinese font support and a game library, plus 4.8MB of supporting files, according to the material. The company said there was no human backstop for testing, code review or manual validation, and that the model even wrote some regression tests and passed them itself.
For the market, the message is that pricing can be tied to completed deliverables rather than interactive minutes—an approach that could reshape how enterprises compare vendors as models begin to behave more like autonomous engineering capacity. Zhipu said its ultimate goal is an Autonomous Agent that can run 24/7, continuously sensing tasks, decomposing goals, executing delivery, self-evaluating, correcting and evolving without human intervention.
Related Coverage:
Zhipu AI Launches Hong Kong IPO Targeting HK$51 Billion Valuation
Zhipu Unveils Flagship AI Model GLM-4.7 Ahead of IPO to Challenge Global Rivals