Why DeepSeek Is Raising Prices While OpenAI Cuts Them

Why DeepSeek Is Raising Prices While OpenAI Cuts Them

What Is Happening?

Something unusual is unfolding in the global AI model market: the pricing curves for American and Chinese large language models are moving in opposite directions.

On one side, OpenAI slashed prices on its GPT-5.6 Luna tier by 80% in late July 2026, and Anthropic quietly abandoned a planned 50% price increase for Claude Sonnet 5. On the other side, DeepSeek announced on its official API pricing page that it would be raising prices across the board — with the increase described as "substantial." Around the same time, Moonshot AI launched its flagship Kimi K3 model in a noticeably higher price bracket: 20 yuan per million input tokens and 100 yuan per million output tokens.

This is not a coincidence or a short-term anomaly. It reflects three structural forces that are simultaneously reshaping the economics of AI model deployment.


Why Are U.S. Models Cutting Prices?

The OpenAI–Anthropic price war

The most visible driver is direct competition between the two dominant U.S. frontier model providers. OpenAI's July price cuts were strategically targeted: Luna and Terra — the tiers that handle heavy agentic workloads — took the biggest reductions. Media reports at the time noted explicitly that the move was designed to pressure Anthropic, whose Claude models hold significant share in enterprise and developer markets but sit at the higher end of the pricing spectrum.

Anthropic's response was telling. Rather than matching the cuts immediately, it canceled a planned price increase for Sonnet 5, its primary agent-use model. The logic is straightforward: in agentic workflows, price directly affects call frequency. A model that costs more per token will be routed around.

The feedback loop that makes low prices strategic

This competition is not purely about revenue. It is about capturing workload data. When an agent runs thousands of real tasks — reading files, writing code, running tests, handling errors — it generates detailed failure signals that inform model improvement. More real-world usage means faster iteration on where agents break down.

The competitive chain looks like this: Lower prices → more real tasks running → more failure signals exposed → better agents → more workload captured

This is why OpenAI has also repeatedly reset usage quotas for Codex and ChatGPT Work subscribers — effectively delivering more compute for the same subscription price. The goal is to reduce the cost of acquiring each unit of real productive workload.


Why Are Chinese Models Raising Prices?

DeepSeek: testing the limits of its own pricing power

DeepSeek has spent the past year building an enormous user base on the back of prices that are, by any measure, extraordinary. Its V4-Flash model costs approximately $0.14 per million input tokens and $0.28 per million output tokens. Independent benchmarking by Artificial Analysis puts the average cost of completing a standard benchmark task at around 3 cents — compared to $1.86 for GPT-5.6 Sol.

Even after a substantial price increase, DeepSeek would almost certainly remain far cheaper than U.S. frontier models. The real question the price increase is designed to answer is not financial — it is strategic: Has adoption become sticky enough to survive higher prices?

If usage holds after prices double or triple, it suggests DeepSeek's global adoption has moved beyond "we use it because it's cheap" into genuine workflow integration. If a large portion of tasks migrate immediately to Qwen, Kimi K3, or other alternatives, that tells a different story — that price was the primary reason for adoption all along.

DeepSeek does not yet have proven pricing power. But it has earned the right to test for it.

Kimi K3: a different strategy entirely

Moonshot AI's approach with Kimi K3 represents a distinct commercial thesis. Rather than competing on price, K3 was positioned from launch as a premium product for long-horizon coding, agentic workflows, and complex knowledge work — and priced accordingly.

The early demand signal was notable: within days of launch, Moonshot temporarily suspended new Kimi subscriptions because demand exceeded available compute capacity, prioritizing existing paying users. That is not proof of durable pricing power, but it is evidence that Chinese open-weight models can enter global markets through a route other than "one-tenth the price of American alternatives."

The internal logic described by people close to Kimi is straightforward: pricing is based on model cost and competitive context within the same use-case segment. DeepSeek, operating in a different task category, is not the reference point.


What Changed the Economics? The Agent Multiplier

Why token price alone no longer captures cost

In the chat-model era, one user query typically triggered one or a few model calls. In agentic workflows, a single user instruction can generate dozens or hundreds of calls: reading files, planning steps, invoking tools, writing code, running tests, reviewing errors, retrying failed steps.

The real cost formula has shifted:

Task cost = Token price × Token volume × Agent steps × Retry rate

This transformation has two important consequences.

First, enterprises now think in terms of cost per successful task rather than cost per million tokens. A cheaper model that requires 40 retries to complete a task may cost more in practice than an expensive model that completes the same task in 8 steps. Independent real-world testing projects like DRadar — which runs models against actual open-source software engineering tasks and measures success rates, latency, and cost simultaneously — show that high-reasoning model tiers achieve better task performance but with costs that rise several times faster than capability.

Second, this economics makes model switching not just possible but rational. Enterprises can now route different steps of the same workflow to different models based on complexity, cost, and required accuracy.


How Enterprises Are Already Responding

The "multi-LLM gateway" pattern

Stanford's Digital Economy Lab published its Enterprise AI Playbook in April 2026, based on 51 successfully deployed AI projects across 41 institutions, 9 industries, and 7 countries. One finding stands out for model providers: 42% of projects considered the underlying model fully replaceable.

For routine, rule-based tasks, that figure rose to 71%. Only for high-stakes decisions requiring complex reasoning did model stickiness increase — and even then, just 35% considered the model a critical differentiator.

Several enterprise operators interviewed for the study described building internal "multi-LLM gateways" that evaluate every incoming request against cost, accuracy, relevance, and latency before routing it to the appropriate model. The operating principle, as one technology company executive put it: "Does this task actually need deep search, or is a mini model sufficient?"

This is the structural pressure that U.S. frontier models now face. Customers have not left — but model loyalty is eroding.

Citi data on open-weight adoption

Data from Citi tracking OpenRouter — a platform used by developers who actively compare and switch between models — showed open-weight models handling 34% of tokens in January 2026 and 65% by June. The four most popular models on the platform during that period were all Chinese. OpenRouter represents a specific developer segment and cannot be extrapolated to the full market, but it illustrates how actively multi-model routing is being practiced among technically sophisticated users.


What Are the Structural Forces at Work?

Three forces are operating simultaneously, and together they explain the pricing divergence:

1. U.S. frontier model competition is pushing high-end prices down as OpenAI and Anthropic fight for agent workload share and the feedback data that comes with it.

2. Chinese open-weight models are continuously expanding the range of tasks that can be completed without premium-priced models. This compresses the set of tasks for which U.S. frontier models can justify high prices.

3. Agentic multi-model routing gives enterprises, for the first time, a practical mechanism to decompose workflows and assign each step to the most cost-effective available model — rather than routing everything through one provider.

The net effect is a price squeeze from both ends. U.S. flagship models are losing the ability to charge large premiums across the full range of tasks. Chinese models that built scale through extreme low pricing are now testing whether they can capture more value per unit of work.


What Comes Next?

The market is not converging to zero — or to permanent premiums

The most likely outcome is not a race to zero where all AI inference becomes a commodity, nor a world where the strongest models maintain 30–50x price premiums indefinitely.

What is being repriced is the task itself — specifically, how much real, recurring, productive work flows through each model or model combination on a sustained basis. Token volume and absolute token price are increasingly poor proxies for commercial value. A model with high call volume may simply be cheap; a model with high prices does not necessarily have pricing power.

The questions that will determine market structure

Several variables will shape how this plays out:

  • Can DeepSeek retain usage after a substantial price increase? The answer will reveal whether its adoption is based on capability or cost.
  • Can Kimi K3 sustain premium pricing at scale? Early demand signals are positive but not yet conclusive.
  • How quickly will enterprise multi-LLM routing mature? As orchestration tooling improves, the ability to substitute models mid-workflow will become easier, further eroding single-model stickiness.
  • Will U.S. frontier models find new moats? The Stanford data suggests complex reasoning, high-stakes decisions, and tasks requiring specialized knowledge remain areas where model choice still matters — and where premiums may hold.

The pricing divergence of mid-2026 is, in this sense, a stress test for every business model in the AI stack. The models that survive it will be those that have built genuine, task-specific value — not just the cheapest option, and not just the most capable one.

Related Coverage:

DeepSeek Signals Major API Price Reset as Demand Tsunami Forces Monetization Reckoning

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe