Morgan Stanley: DeepSeek’s V4 Flash Price Cut Still Leaves It Behind Cheaper Rivals
DeepSeek's headline price cut on its V4 Flash model signals competitive stress rather than market dominance, with Morgan Stanley concluding the move fails to close a widening pricing gap against cheaper Chinese rivals.
Effective September 10, 2026, DeepSeek slashed pricing across three tiers of its V4 Flash model — non-peak cached input fell 60%, non-cached input dropped 33%, and output declined 11% — bringing non-cached input to RMB 1.5 per million tokens (approximately US$0.21), cached input to RMB 0.03, and output to RMB 6. The announcement, made on September 9, triggered immediate scrutiny from institutional desks tracking China's accelerating large language model (LLM) price war.
Morgan Stanley analysts Gary Yu, Lydia Lin, and Yang Liu identified two concurrent catalysts behind the timing: competitive encroachment from models launched in August 2026, and a strategic clearance move ahead of DeepSeek's own V4.1 Flash release. The bank's assessment, however, carries a cautionary undertone — the price reduction, while substantial in percentage terms, does not resolve DeepSeek's structural pricing disadvantage in the lightweight segment.
Competing Models Force DeepSeek's Hand
August 2026 saw two direct challengers enter the lightweight LLM tier simultaneously. Alibaba launched Qwen3.8-flash, priced at RMB 1 per million tokens for non-cached input and RMB 3 for output. Zhipu AI moved more aggressively with GLM5.3-flash, setting non-cached input at RMB 0.4 and output at just RMB 1.4 — the latter representing less than one-quarter of DeepSeek's post-cut output price.
According to Morgan Stanley's comparative data, the performance differential across the three models does not justify DeepSeek's premium. On the AA Intelligence Index, GLM5.3-flash scores 57 points, Qwen3.8-flash 56, and DeepSeek V4 Flash 52 — a gap that undercuts any quality-based pricing rationale. DeepSeek V4 Flash carries 284 billion parameters versus Qwen3.8-flash's 125 billion and GLM5.3-flash's 320 billion, but raw parameter count has demonstrably failed to translate into benchmark leadership or pricing power.
Post-Cut Pricing Still Exceeds Launch Levels, Exposing a Structural Overhang
Morgan Stanley's analysis surfaces a detail that complicates DeepSeek's narrative: even after the cuts, V4 Flash remains more expensive than its own April 2026 launch price. At debut, non-cached input was RMB 1 per million tokens and output was RMB 2. The current post-cut figures of RMB 1.5 and RMB 6, respectively, represent a net price increase of 50% on input and 200% on output relative to initial positioning — meaning the September reduction is a partial reversal of prior hikes, not a genuine market-share offensive.
Against competitors, the gap is stark. DeepSeek's output price of RMB 6 is double Qwen3.8-flash's RMB 3 and more than four times GLM5.3-flash's RMB 1.4. On non-cached input, DeepSeek at RMB 1.5 sits 50% above Alibaba's offering and nearly four times Zhipu AI's rate. Morgan Stanley's conclusion is direct: the price adjustment does not alter DeepSeek's position as the highest-cost option among mainstream lightweight models.
Pro-Plus-Flash Dual-Track Strategy Becomes Industry Standard
Morgan Stanley frames DeepSeek's move within a broader structural shift it identifies across China's LLM vendor landscape. The bank argues that a bifurcated product architecture — pairing high-capability, high-cost "Pro" models targeting enterprise precision workflows with low-latency, low-price "Flash" variants aimed at high-frequency developer and consumer use cases — is now the de facto competitive template.
This dual-track model mirrors dynamics seen in cloud infrastructure pricing, where tiered offerings allow vendors to defend margin on premium SKUs while using entry-level products to capture volume and ecosystem lock-in. As competition intensifies specifically at the Flash tier, Morgan Stanley anticipates sustained pricing pressure across that segment through the remainder of 2026.
Gross Margin Constraints Seen Capping Price War Escalation
Despite the competitive intensity, Morgan Stanley pushes back against the most bearish scenario — a race to zero. The bank's analysts argue that gross margin constraints will function as a natural floor, preventing vendors from sustaining below-cost pricing strategies. Under profitability pressure, the bank expects LLM providers to compete on differentiated performance benchmarks and ecosystem integration rather than pure token-price undercutting.
This view carries implications for investors monitoring China's AI infrastructure spending cycle. If margin discipline holds, the current pricing dynamic may stabilize around existing tiers rather than compressing toward commodity levels — a scenario that would favor vendors with diversified monetization channels over pure-play model providers dependent on API revenue.
Related Coverage:
DeepSeek Cuts Flash API Prices by 60% as the Fight for Developers Intensifies