Cloud Giants Push Through AI-Driven Price Hikes as Token Demand Strains Compute and Networks
Global cloud providers are resetting pricing power in 2026 as AI inference demand collides with higher supply-chain and infrastructure costs, accelerating a broad-based shift from “scale-first” growth to margin protection.
Alibaba Cloud and Baidu AI Cloud on March 18 became the latest major vendors to raise prices for AI compute and storage products, extending a wave of increases that has already touched Amazon AWS, Google Cloud and Tencent Cloud. The immediate takeaway for investors and enterprise buyers is clear: AI workloads are no longer deflationary at the infrastructure layer, and network-heavy services are emerging as the next battleground.
The repricing is not confined to China. Google Cloud has lifted prices on several data transfer services, including CDN Interconnect, Direct Peering and Carrier Peering, with North America hikes reaching 100%. Tencent Cloud’s large-model service adjustment has been even sharper in select cases, with the input price for Tencent HY2.0 Instruct rising 463.13% to RMB 0.004505 per 1,000 tokens from RMB 0.0008.
Alibaba Cloud and Baidu AI Cloud lift AI infrastructure pricing
Alibaba Cloud said it will raise prices for AI compute, storage and related products due to a global surge in AI demand and rising supply-chain costs. The company set a maximum increase of 34%, with price changes of 5% to 34% for compute card products such as T-Head Zhenwu 810E, and a 30% rise for CPFS (Intelligent Computing Edition) file storage.
Baidu AI Cloud announced a “structural optimization” of pricing for parts of its AI compute and storage portfolio, citing rapid growth in AI applications, sustained increases in compute demand and “significant” rises in core hardware and infrastructure costs. The company framed the move as necessary to maintain long-term platform stability and service quality.
Together, the same-day moves signal that China’s leading hyperscalers are aligning more closely with global peers: pricing is being recalibrated around AI-era unit economics, rather than legacy cloud price-war dynamics.
Network and data transfer costs broaden the repricing beyond GPUs
Price adjustments are increasingly concentrating in “data transfer and network” categories, where increases are commonly in the 10% to 40% range, according to the material. AWS, Google Cloud, Microsoft Azure, Tencent Cloud and Wangsu Science & Technology have all included network-related services in their price actions, indicating vendors are passing through rising bandwidth and network infrastructure costs.
Google Cloud’s 100% increase in specific North American network services stands out as an extreme example, but the cadence is equally notable: major vendors have been announcing hikes almost monthly, creating a follow-the-leader effect across the industry.
Wangsu product director Wang Zhijie said the cloud price-war phase has ended and the market is entering a “value reversion” cycle, shifting from “scale priority” to “profit priority.” He added that network transmission is a cloud provider’s second-largest cost item after compute, suggesting CDN pricing is structurally re-rating as AI workloads demand low-latency connectivity between edge nodes and central cloud.
Inference demand and agent workloads tighten capacity and change unit economics
The current wave is being driven less by training and more by inference. Wang Zhijie observed that from 2025 through the first quarter of 2026, training demand has been relatively stable while inference demand has grown exponentially. Industry data cited in the material show large-model API calls rising about 30% month-on-month, while video generation and real-time interactive applications are pushing up edge inference compute needs.
A cloud industry practitioner described a key break from traditional cloud economics: unlike the classic “Moore’s law plus scale effect” cost-down path, AI compute’s marginal cost can rise with scale, creating a risk that “the more you sell, the more you lose.” That dynamic makes structural price increases a tool to repair margins as capacity tightens across GPUs, storage, bandwidth and power.
The same practitioner also highlighted operational complexity as a cost driver: the challenge is not only securing resources, but scheduling heterogeneous compute (CPU, GPU, FPGA), supporting seamless migration across edge, core and cloud, reducing edge model loading latency, and upgrading liquid cooling and power as rack-level power density rises.
Token surges from OpenClaw and agents amplify pricing pressure
AI agent workflows are emerging as the multiplier behind the infrastructure shock. A person familiar with the matter said Alibaba Cloud’s MaaS platform Bailian posted its highest growth rate ever from January to March 2026, prompting the company to tilt scarce AI capacity toward token-based inference services as token calls surged.
OpenRouter data cited in the material show OpenClaw token consumption jumping from 80.6 billion on Feb. 3, 2026 to 358.0 billion by March 4, a roughly 3.4x increase in about one month. For the week of March 2, OpenRouter’s weekly token calls reached 14.8 trillion, up about 160% in two months, with OpenClaw contributing most of the total. Anthropic’s data indicate agent workloads can consume up to 15 times the tokens of ordinary chat interactions.
At Nvidia’s GTC on March 17, CEO Jensen Huang said AI agents often require repeated inference calls across multiple models and tools for a single task, driving token consumption up by orders of magnitude. Huatai Securities said faster rollout of “Claw-like” products could accelerate agent evolution and keep token consumption, inference compute demand and infrastructure investment trending higher.
IDC projected token consumption will surge as agents become more complex, with annual usage expected to rise from 0.0005 Peta Tokens in 2025 to 152,667 Peta Tokens by 2030, implying a CAGR of 3,418%. IDC China senior research manager Sun Zhenya said costs and energy consumption will be key constraints, advising enterprises to plan ahead on compute procurement, model selection and model combinations.
Related Coverage:
AI Cloud Race Intensifies as Alibaba, ByteDance Compete for Market Leadership