DeepSeek Designs AI Inference Chip to Cut Nvidia, Huawei Reliance
DeepSeek, the Hangzhou-based AI lab whose cost-efficient large language models rattled global semiconductor markets earlier this year, is developing a proprietary AI inference chip in a direct bid to reduce structural dependence on both Nvidia and Huawei Technologies, according to three people with knowledge of the matter cited by Reuters on July 7, 2026.
The disclosure marks a meaningful strategic inflection point for a company that has until now competed almost entirely on software and algorithmic efficiency. By moving into silicon, DeepSeek is signaling that the compute constraint—not model architecture—has become its binding bottleneck, and that founder and CEO Liang Wenfeng views hardware self-sufficiency as a non-negotiable long-term condition for the lab's independence.
DeepSeek declined to comment when approached by Reuters. The company did not respond to additional requests for comment from this publication.
Export Controls Force DeepSeek to Accelerate Hardware Strategy
The chip development effort began approximately one year ago—placing its origins in mid-2025—and remains in early-stage development, the sources said. DeepSeek is currently in active discussions with chip design firms, contract wafer foundries, and memory suppliers as it maps out an external partnership ecosystem to complement its internal engineering work.
The timing is not incidental. In late 2023, the U.S. Commerce Department banned exports of Nvidia's H800 GPU—a chip that had served as the primary training substrate for several of DeepSeek's flagship models—to Chinese entities. That single regulatory action compressed DeepSeek's hardware optionality and forced a rapid pivot toward Huawei's Ascend series accelerators.
In April 2026, DeepSeek released a version of its V4 model optimized for Huawei's Ascend chips. Huawei confirmed that Ascend silicon contributed to portions of the V4-Flash lightweight model's training process. Reuters separately reported that the V4 release triggered a material surge in Chinese technology companies' orders for Huawei's Ascend 910C chip—an indirect but measurable validation of DeepSeek's influence on domestic chip demand.
The strategic logic is straightforward: DeepSeek currently operates a dual-vendor hardware stack, running both Nvidia and Huawei chips in parallel. Each vendor carries distinct geopolitical risk. A proprietary inference chip—even one that supplements rather than replaces third-party silicon—would give DeepSeek a hardware layer it fully controls, reducing exposure to either Washington's export regime or Beijing's industrial policy shifts.
Liang Wenfeng's Compute Philosophy Foreshadowed the Pivot
Liang's public statements over the past three years reveal a consistent preoccupation with compute scarcity that, in retrospect, reads as strategic groundwork. In two separate interviews with Chinese media outlet Anwave conducted in 2023 and 2024, Liang stated: "Our real challenge has never been capital—it has always been the export ban on high-end chips," and separately, "For researchers, the appetite for compute is infinite. We will deliberately deploy as much compute as possible."
Those remarks, made before DeepSeek's global breakthrough, now carry the weight of a founding thesis. A lab that identified chip access—not funding—as its primary constraint three years ago was, by definition, already thinking about the conditions under which it might need to own its supply chain.
The recruitment pattern reinforces this reading. Sources told Reuters that DeepSeek has significantly increased hiring of chip design engineers in recent months. Critically, these positions have not been posted on public platforms such as BOSS Zhipin or Liepin; recruitment is being conducted through private, direct-sourcing channels—a deliberate approach that limits competitive intelligence leakage and suggests the program carries strategic sensitivity at the executive level.
In-House Silicon Becomes the New Table Stakes for Frontier AI Labs
DeepSeek's move aligns it with a converging global consensus among top-tier AI developers that inference-optimized, proprietary silicon is a competitive necessity rather than a luxury.
OpenAI has partnered with Broadcom to develop its first custom inference chip. Anthropic is reportedly evaluating a comparable in-house chip program. In each case, the strategic rationale is consistent: full replacement of third-party silicon is not the objective. Rather, hardware-software co-optimization—tuning silicon specifically to a lab's model architectures and inference workloads—offers measurable gains in energy efficiency, per-token cost reduction, and supply chain control that general-purpose GPUs cannot match at scale.
For DeepSeek, the efficiency imperative is especially acute. The lab built its reputation on achieving frontier-level model performance at dramatically lower compute costs than Western peers—a differentiation that is inherently hardware-sensitive. A purpose-built inference chip would allow DeepSeek to extend that cost advantage into the deployment layer, potentially widening the margin between its inference economics and those of competitors running on commodity accelerators.
V4 Full Release Imminent, Pricing Model Shifts to Peak-Valley Structure
The chip development news coincides with DeepSeek's most significant near-term product catalyst. The company notified API customers via email last week that the full commercial release of DeepSeek V4 is scheduled for mid-July 2026. The release will introduce a peak-valley pricing structure: API rates will double during peak usage hours while remaining unchanged during off-peak periods.
Early signals from the Chinese user community suggest the full V4 release is already in limited grey-scale testing. Users reporting access to what appears to be the production build have flagged a material improvement in code generation capabilities—a benchmark category that carries outsized weight with enterprise API customers.
The pricing architecture shift is itself analytically significant. A move from flat-rate to time-differentiated API pricing indicates that DeepSeek is experiencing genuine demand-side capacity pressure—a constraint that a proprietary, inference-optimized chip is precisely designed to address.
Related Coverage:
OpenClaw Crowns DeepSeek-V4 as Default AI Model Amid Integration Turbulence