Morgan Stanley: DeepSeek’s "Engram" May End the AI Hardware Arms Race

Morgan Stanley: DeepSeek’s "Engram" May End the AI Hardware Arms Race

In the escalating artificial intelligence war between the U.S. and China, the prevailing narrative has been simple: he who possesses the most Nvidia H100s (or H200s) wins. The U.S. strategy of choking off China's access to advanced High Bandwidth Memory (HBM) and cutting-edge GPUs was designed to cripple the Middle Kingdom's AI ambitions.

However, a new research note from Morgan Stanley suggests this strategy may have inadvertently birthed a far more dangerous competitor: one that doesn't need brute force to compete.

In a report dated January 21, 2026, analysts led by Shawn Kim and Charlie Chan highlight a pivotal shift in AI architecture driven by Chinese AI unicorn DeepSeek. The bank argues that DeepSeek’s latest innovation, the "Engram" module, effectively decouples memory from compute, allowing Chinese models to bypass hardware bottlenecks and deliver frontier-level performance on consumer-grade chips.

Constraint-Induced Innovation

While Silicon Valley hyperscalers continue to burn billions on massive GPU clusters, DeepSeek has been forced to innovate out of necessity. Morgan Stanley terms this dynamic "constraint-induced innovation."

The core problem for China has been the scarcity of HBM—the ultra-fast memory located next to the GPU processor, which is essential for running massive Large Language Models (LLMs). DeepSeek’s solution, detailed in their January 12, 2026 paper, is a move toward a hybrid architecture that relies less on expensive HBM and more on cheap, abundant system memory (DRAM).

Morgan Stanley explains the technical breakthrough:

"DeepSeek's latest innovation of Engram module reduces HBM constraints and infrastructure costs via decoupling storage from compute... Engram is an approach to help ease memory constraints across AI infrastructure and efficiently 'look up' essential information without overloading HBM, thereby freeing capacity for more complex reasoning tasks."

Essentially, instead of forcing the GPU to hold every piece of information in its expensive, fast memory, Engram creates a "library" of static information stored in standard DRAM. The model then performs a low-cost "lookup" only when necessary.

Smarter, Not Just Bigger

The implications of this architectural shift are profound. For the last three years, the industry has operated under the assumption that "bigger is better." DeepSeek is proving that "smarter is cheaper."

The report notes that by offloading memory to a scalable lookup system, DeepSeek can achieve reasoning improvements that exceed knowledge gains. This suggests that the value in the AI stack is shifting from raw compute to efficient memory management.

Most strikingly, Morgan Stanley predicts that DeepSeek’s upcoming V4 model could run effectively on consumer hardware.

"DeepSeek's next generation LLM V4 could be a major leap forward when released... utilizing Engram memory architecture. Like its predecessor, this model most likely can run on consumer hardware with consumer-grade hardware (RTX 5090) potentially sufficient."

If an RTX 5090 can handle inference tasks that previously required data-center class chips, the unit economics of AI collapse in China's favor.

Closing the Gap

Despite operating under severe hardware sanctions, Chinese models are rapidly closing the performance gap with U.S. frontier models like ChatGPT 5.2.

The report highlights that models such as DeepSeek-V3.2 and Qwen-3 (developed by Alibaba) are achieving comparable scores on standardized benchmarks like MMLU and GPQA. They are doing this at a fraction of the compute cost, utilizing sparse Mixture-of-Experts (MoE) architectures and the new Engram memory offloading.

The bank notes: "China’s AI development is increasingly shaped by a 'constraint-induced innovation' dynamic... In the long run, this may produce AI ecosystems that are more cost-effective, scalable, and adaptable."

Investment Implications: The Rise of Local Semi-Caps

Morgan Stanley believes this architectural pivot reinforces the bull case for China's domestic semiconductor supply chain. If the future is hybrid architecture and efficient memory usage, the companies building the tools for this ecosystem stand to benefit.

The bank maintains an "Overweight" rating on key players in the localization theme, specifically highlighting:

  • NAURA Technology Group Co Ltd : Target price RMB 514.2 (US$71.41).
  • Advanced Micro-Fabrication Equipment Inc : Target price RMB 364.32 (US$50.60).
  • JCET Group Co Ltd: Target price RMB 49.49 (US$6.87).

The logic is straightforward: As China builds LLMs that are structurally more efficient, the spending momentum for domestic equipment and memory solutions will accelerate.

The Bottom Line

The U.S. sanctions regime was predicated on the idea that hardware superiority is the only path to AI dominance. DeepSeek’s Engram suggests that software efficiency can effectively arbitrage hardware deficits.

By forcing China to "think around constraints," the West may have accelerated the development of a more efficient, lower-cost AI architecture that could eventually undercut the massive, power-hungry models being built in Silicon Valley. As Morgan Stanley concludes, the next AI frontier "may not be simply bigger models but efficient hybrid architectures."

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe