China's Smartphone Giants Pivot to On-Device AI in Race for Dominance
Chinese smartphone manufacturers are pivoting away from large-scale foundation models to concentrate on lightweight, on-device multimodal models, seeking to deliver advanced AI capabilities directly on handsets while managing costs and privacy concerns. Vivo, OPPO, and Honor unveiled their latest AI strategies at recent developer conferences, showcasing capabilities ranging from screen recognition to automated bill recording, marking a significant evolution from text-focused applications that dominated two years ago.
The shift reflects a maturing approach to AI integration, with companies deploying 3B-parameter models on devices that match the performance of much larger cloud-based systems. This transition addresses mounting operational costs—cloud-based voice recognition alone can cost RMB 2 yuan (US$0.28) per hour of use—while addressing user privacy concerns that limit cloud model data access.
However, the technology's advancement faces a critical bottleneck: the absence of breakthrough applications that could justify higher chip costs and drive broader adoption. Chipmakers including Qualcomm and MediaTek, whose latest flagship processors deliver 100 TOPS of AI computing power, are pressing manufacturers to demonstrate compelling use cases that warrant premium pricing.
The emerging agent ecosystem, while promising cross-app automation, remains largely confined to manufacturers' native applications, with third-party integration hampered by unresolved questions around data sharing and traffic distribution.
On-Device Multimodal Models Take Center Stage
Smartphone AI has progressed substantially from text processing tasks like dialogue generation and summarization that relied on cloud computing. This year's primary innovation involves deploying multimodal on-device models capable of handling image and voice processing locally.
Vivo demonstrated 18 on-device intelligent applications, including ID card recognition, automatic file naming, and on-device UI agents that create notes or record detailed bills through single voice commands. These tasks involve more complex interaction logic than previous simple functions, requiring intent recognition and autonomous planning capabilities.
OPPO highlighted its one-tap screen query and one-tap flash memo features. The screen query function, powered by multimodal models, enables real-time screen content understanding and allows users to point at real-world objects for voice interaction. The flash memo automatically extracts and categorizes key information—recording WeChat payment transactions to bills without manual input, scanning receipt data for bill entry, or displaying pickup codes as real-time reminder cards.
Honor disclosed that its phones can automatically execute over 3,000 scenarios across daily activities. These include one-tap price comparison shopping that adds items to carts and applies coupons, and one-tap ride-hailing through voice commands. Tasks previously requiring frequent app switching now complete through single AI instructions.
"From popular large models and agent products, the technology has developed the capability to understand the physical world or accelerate the integration of physical and digital worlds," said Zhang Chong, general manager of Honor's MagicOS AI product division. The digital world contains natural data and production data that enables better model fine-tuning to understand user needs in current contexts.
An AI technology expert at a smartphone company noted a mismatch between AI progress and user needs: "Users' most frequent AI use case is image processing, but this generation's technology matured first in language models." The expert predicted image processing maturity will reach very high levels by next year.
Technical Evolution Toward Lightweight Models
Smartphone manufacturers' large model development has progressed through three stages. Two years ago, Vivo and OPPO released full-size language models ranging from hundreds of millions to hundreds of billions of parameters. A year ago, industry focus shifted from language models toward multimodal domains including voice and images, with increased emphasis on on-device deployment.
This year shows several clear trends. First, on-device models concentrate on lightweight 3B parameter sizes while adding multimodal capabilities to language models. In July, Honor released its 7B multimodal perception model MagicGUI. Vivo simultaneously launched its 3B multimodal reasoning model BlueLM-2.5-3B, integrating language, vision, and logical reasoning on-device. In October, OPPO released its on-device multimodal model AndesVL, featuring four size configurations from 0.6B to 4B with general multimodal recognition, understanding, reasoning, GUI capabilities, and multilingual support.
The industry has rapidly reduced model size and memory overhead through low-bit mixed quantization and on-device LoRA training solutions, accelerating on-device multimodal model deployment. Current 3B models achieve the performance of previous 8B models, according to an industry practitioner. Tasks previously requiring multiple visual expert models combined with language models now integrate multiple sizes and modalities into single models, delivering higher recognition rates. Vivo employs a 1+N architecture where multimodal, language, and logical reasoning models share a base model, paired with modal-specific LoRAs to support over ten business scenarios.
Second, on-device reasoning models now implement deep thinking modes, enabling phones to perform complex cloud-level reasoning locally with significantly improved accuracy for complex problems.
Third, GUI Agent model introduction enables AI to actively control phone interfaces for task completion. It simulates human clicking and swiping operations without relying on rules, fixed scripts, or special APIs from applications, allowing phone agents to operate third-party apps.
Deployment Challenges and Cost Pressures
Current phone AI assistants typically invoke different models for different tasks, including proprietary distilled models and external cloud-based services via APIs. Alibaba's Tongyi and ByteDance's Doubao are widely integrated by phone manufacturers.
However, a smartphone industry source revealed complications in calling external models. "Whether Doubao or Alibaba, the APIs they provide phone manufacturers differ from their latest internal versions, lagging at least three to six months," the source said. Cloud vendors internally separate teams selling cloud services from those developing models.
Cloud vendors package internal capabilities as commercial products, but model makers worry phone manufacturers might optimize these with proprietary data to achieve superior results. "It's not that I don't want to integrate them, it's that they don't want to provide to me," the source said.
Compared to two years ago, phone manufacturers invest less in large-parameter foundation models, focusing instead on on-device multimodal models. A phone AI expert explained that while cloud models achieve substantial compression through MOE architecture, on-device models constrained by chip performance now reach 2B-5B parameters, equivalent to 2023's 32-70B models. Model vendors pursue maximum intelligence, while terminal manufacturers focus on compression for on-device deployment. "We don't do zero-to-one foundation model training; small on-device models are actually distillations of large cloud models."
"Cloud-side capabilities are relatively easy to establish," said Zhou Wei, director of Vivo's AI Research Institute. "What's truly difficult is on-device capability."
Zhou disclosed that Vivo developed 13B and 7B on-device models last year, finding only 7B basically usable, but with unsatisfactory performance requiring nearly 4GB RAM. Over the past year, Vivo concentrated efforts on 3B on-device multimodal models. Current 3B models achieve 97%-98% of cloud model capability for text summarization—"already sufficient."
This doesn't mean phone manufacturers abandon large-parameter models, but rather differentiate capabilities. "If a problem is already being solved by most manufacturers, then I choose to cooperate with them," a technical expert said. Phone makers won't iterate models purely adding world knowledge, but focus on understanding multi-dimensional phone data to pursue personalized intelligence.
While manufacturers currently employ cloud-device collaboration, focus remains on on-device model optimization. Cloud API calls incur per-use costs with latency affecting user experience, while privacy concerns limit cloud model data usage. On-device models, beyond requiring higher-performance chips and storage, add virtually no other costs while offering superior privacy through local processing—key advantages for phone deployment.
AI's explosion brings phone manufacturers sweet burdens. Large user bases making frequent cloud service calls generate enormous costs. Using ASR models for phone transcription translation costs RMB 2 yuan per hour, expenses hardware makers must absorb, according to a phone AI expert.
A practitioner noted cloud vendors lack strong incentives for on-device model investment "because they primarily sell MaaS services," increasing reliance on phone manufacturers to solve on-device model challenges independently.
Current issues include the absence of breakthrough AI applications, limited user AI awareness, and resulting hesitation among chipmakers. "Chip manufacturers keep approaching us to find more flagship scenarios on phones," the source said. Qualcomm Snapdragon and MediaTek Dimensity latest flagship chips already deliver 100 TOPS AI computing power. Chip vendors want to sell higher-powered chips, but without sufficient application support, greater computing power means higher chip prices, ultimately affecting sales.
Agent Ecosystem in Early Stages
Currently visible one-command photo editing, Wi-Fi connection, and bill recording automated tasks remain largely confined to manufacturers' native apps like notepads and photo galleries.
However, users spend most time in third-party applications—"85% of duration comes from services provided by developers"—meaning participation from leading internet companies remains crucial.
Zhou noted that current phone autonomous agents executing tasks can only handle manufacturers' own functions. Cross-app capabilities require complex discussions between terminal and internet manufacturers around security authorization standards. "As terminal manufacturers, we must actively promote industry standard establishment while recognizing that AI technology requires several more years to mature."
As single agents evolve toward multi-agent coordination, phone manufacturers besides releasing agent applications actively build agent ecosystems.
Vivo extracts high-frequency reusable system capabilities into general system-level agents, including screen perception and task planning as "general control facility groups" for ecosystem partners to invoke directly. Through its agent development platform providing multiple on-device AI development capabilities, Vivo helps ecosystem partners develop diverse agents for specific business scenarios.
OPPO positions its agent ecosystem framework as one of three technical cornerstones of OPPO AI—not only the core platform for OPPO agent cross-device coordination but key to upgrading AI agents from single-step execution to complex task planning and multi-device collaboration.
Honor released its system-level MCP architecture, currently connecting over 80% of high-frequency system-level scenarios and integrating over 4,000 ecosystem MCPs and agents. Beyond software ecosystems, Honor leverages Shenzhen's location advantages to build AI hardware ecosystems enabling agent cross-device coordination.
Phone manufacturers' agent ecosystem construction enjoys natural advantages over other terminal products through abundant cross-app, cross-scenario multimodal data. Phones connecting with other devices can serve as intelligent hubs—characteristics providing inherent advantages in agent ecosystem building.
Internet companies are beginning to benefit. Ant Group has established strategic cooperation with mainstream phone manufacturers, integrating its agent services into phone ecosystems. Vivo disclosed that Ant's AI health agent AQ's traffic share in Blue Heart Xiao V's health scenarios tripled from early this year to now.
For most application vendors, agent ecosystems involve difficult questions around traffic distribution and data permissions. Many app companies worry system-level agents directly serving end users could undermine app value. Additionally, whether user data currently controlled by individual apps needs sharing for system-level agent execution concerns many enterprises.
The industry's current common approach develops GUI models as a gentler solution—essentially not agent-to-agent interaction, but AI replacing human interface operations. Users still log into personal accounts with AI confirmation required at key points, with phone agents playing user roles.
Vivo's Zhou Wei's stance represents many phone manufacturers' views: "First, we'll sit down and discuss with those willing to shake hands. Second, with AI's arrival, whether there needs to be a completely new pecking order and influence, we'll leave that to time."