DeepSeek V4.1 Flash: The Architecture Shift Redefining AI Inference Economics

DeepSeek V4.1 Flash: The Architecture Shift Redefining AI Inference Economics

DeepSeek released V4.1 Flash on Thursday, making the model available simultaneously across its web interface, mobile app, and API. The release marks a significant architectural departure from the previous generation and comes with a pricing structure that undercuts its predecessor.

Architecture and Design

V4.1 Flash is a Mixture-of-Experts model with 552 billion total parameters, but its defining feature is an asymmetric computation design: the input side activates only 8 billion parameters, while the output side activates 16 billion. DeepSeek describes this as a Causal-Encoder-Decoder architecture — a new framework for the company — intended to raise capability ceilings while lowering per-inference costs and improving throughput. The company positions V4.1 Flash as the smallest model in this new architectural series, with larger variants implied for future releases.

The model was trained using a new pre-training methodology and a larger-scale reinforcement learning post-training pipeline. DeepSeek has open-sourced the model weights on Hugging Face alongside a technical report.

Benchmark Performance

Benchmark results show uneven but notable progress. Compared with V4 Pro, V4.1 Flash improved significantly on agent-oriented tasks: DeepSWE v1.1 rose from 62.7 to 74.2, Terminal-Bench 3.0 from 11.8 to 30.0, Automation-Bench from 43.2 to 54.8, and CyberGym from 83.3 to 88.1. On DeepSWE v1.1, CyberGym, and Automation-Bench, V4.1 Flash posted the highest scores among compared models, which included Kimi K3, GLM 5.3, Claude Opus 5, and GPT-5.6 Sol.

However, the gains are not uniform. V4.1 Flash trails Opus 5 on Terminal-Bench 3.0 — scoring 30.0 versus Opus 5's 43.3 and GPT-5.6 Sol's 34.4 — and still lags behind leading closed-source models on GPQA Diamond and ProgramBench. The overall picture suggests a model competitive in coding and agentic workloads, but not uniformly dominant across all benchmarks.

KV Cache Compression and Cost Reduction

One of the more technically consequential changes is a dramatic reduction in KV Cache size. According to DeepSeek, V4.1 Flash reduces per-token KV Cache from 3,514 bytes in the previous V4 Flash to 890 bytes — approximately one-quarter of the prior generation and roughly 1/437th of the original DeepSeek V1. The company states this translates to HBM memory requirements falling to one-quarter and SSD storage requirements to one-eighth compared with the previous generation.

The practical consequence is that weight reads — not cache reads — now dominate memory bandwidth during inference. DeepSeek says this is why the model sustains generation speeds of around 400 tokens per second without degradation as context length grows, a characteristic particularly valuable for long-context and multi-turn agentic tasks where input tokens vastly outnumber output tokens.

Pricing

New API pricing took effect at 12:00 Beijing time on September 10, 2026. DeepSeek uses a peak/off-peak pricing structure, with off-peak rates set at half the peak price. All figures are denominated in Chinese yuan per million tokens.

During off-peak hours: cache-hit input costs RMB 0.02 yuan (approximately US$0.003), cache-miss input costs RMB 1 yuan (approximately US$0.14), and output costs RMB 4 yuan. Peak-hour rates double across all categories. Peak hours are defined as weekday Beijing time 9:00–12:00 and 14:00–18:00; all other times, including weekends, are off-peak.

The 50-fold difference between cache-hit and cache-miss input pricing creates strong economic incentives for developers to structure workloads around prefix reuse — a design choice aligned with the model's architectural strengths in agent and long-context tasks.

V4 Pro Deprecation

DeepSeek has announced that V4 Pro's API will be taken offline at 12:00 Beijing time on September 14, 2026. From that point until the future launch of V4.1 Pro, all requests routed to the deepseek-v4-pro endpoint will be automatically redirected to V4.1 Flash and billed at V4.1 Flash rates. The company's decision to deprecate its previous flagship model within days of the new release signals confidence in V4.1 Flash as a direct replacement.

On the consumer side, DeepSeek has consolidated its web and app interface: the previous Fast, Expert, and Vision modes have been merged into a single entry point, all powered by V4.1 Flash.

Harness Update and Ecosystem

DeepSeek Harness, the company's agentic execution environment, has been updated to version 0.1.5 in tandem with the model release. The new version was trained specifically against V4.1 Flash across standard, Programmatic Tool Calling, and minimal modes. New features include file upload support for images and PDFs, a sidebar file preview panel, bidirectional communication between parent and child agents during task execution, and an experimental Agent Teams function that allows a primary agent to delegate work across multiple sub-agents via a shared task list. The Agent Teams feature is disabled by default and requires opt-in through an experimental plugin.

Tencent's WorkBuddy and CodeBuddy, along with OpenCode, have been named as official partners and have integrated the new model. Third-party platform Workbuddy has also announced support. For teams with large-scale deployment needs and approximately 2,000 GPUs and storage cluster resources, DeepSeek has opened a channel for deployment partnerships.

Related Coverage:

DeepSeek Prepares for a STAR Market IPO as Its Valuation Nears $70B

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe