Tech Buzz China Insider

Tech Buzz China Insider

Why Everyone in China Wants an AI Inference Chip

Restricted Nvidia access opened the market. The economics of serving AI are pushing chipmakers, cloud providers and model developers toward custom hardware.

Tech Buzz China's avatar
Tech Buzz China
Jul 13, 2026
∙ Paid

Things you might have missed

A few items from the past week that didn’t make it into the main analysis but are worth tracking.

From our recent coverage:

  • Unitree’s vertical integration extends to raw materials. Co-founder Chen Li said the company’s supply chain for joint motors extends little beyond copper wire and magnets, reflecting a strategy of building nearly every core component in-house. Read more →

  • Alibaba reverses its 2023 breakup strategy. In one week, it folded Cainiao’s domestic supply chain back into its China commerce group and merged its Spain-focused local e-commerce operation into AliExpress, consolidating power under Jiang Fan. Read more →

  • Tencent’s Hunyuan 3 ships under new AI chief. The official release of Hy3 is the first model built from scratch under Yao Shunyu, who joined last December and immediately rebuilt the pretraining and reinforcement learning infrastructure. The 295B-parameter MoE model activates only 21B per query. Read more →

From Weijin Research:

  • Dedicated inference hardware is splitting up AI compute. Our partner publication also examined how the shift from training-optimized GPUs to specialized inference chips is rewriting priorities from FLOPS to latency, efficiency, and memory bottlenecks. Read the full analysis →

Where we have been featured:

  • CNBC (July 11, 2026): While Musk’s Neuralink drills into skulls, China’s BrainCo bets the future of brain tech is wearable.

  • Semafor (July 7, 2026): “FOBO” is driving China’s AI anxiety.

China’s inference market did not emerge overnight. In March, we examined how DeepSeek rewrote the AI playbook through model efficiency and cost discipline. Last week’s piece on the month China closed the AI stack showed model developers, chipmakers and cloud providers converging around domestic infrastructure. In September, we explored Cambricon’s valuation paradox—a trillion-renminbi market capitalization supported more by future potential than current shipments. This piece is the next step in that thread.

The Inference Rush

Within hours of each other last week, Reuters reported that DeepSeek was exploring an inference chip, while The Information reported that Z.ai had begun preliminary discussions with domestic chip-design firms about building a processor optimized for its GLM models. Neither company has announced a partner, architecture or production timetable, and both projects remain at an early stage.

DeepSeek is building its own AI chip. Reuters reports they are designing an inference  chip to cut its dependence on Nvidia and Huawei. It's already in talks with  chip-design, foundry, and memory

On its face, the announcements are surprising. DeepSeek already serves inference on Huawei’s Ascend processors, while Z.ai has spent the past year adapting GLM to processors from Huawei, Cambricon, Moore Threads and Kunlunxin. If DeepSeek already runs on Huawei, why build another processor?

The question extends well beyond two companies. Across China’s AI industry, chipmakers are commanding extraordinary valuations, internet platforms are spinning out semiconductor businesses and model developers are moving steadily closer to the hardware they run on.

On June 30, Cambricon became China’s first trillion-renminbi AI chip company. Its market capitalization reached roughly $138 billion, or about 373 times trailing earnings, even though IDC estimates it shipped only 116,000 accelerator cards in 2025, equal to 2.9% of the domestic market. Investors were valuing the market Cambricon might someday serve more than the one it serves today.

Internet platforms are making the same bet. Alibaba is moving toward an independent listing of T-Head after deploying more than 100,000 self-developed Zhenwu accelerators across Alibaba Cloud, while Baidu’s Kunlunxin is pursuing a Hong Kong IPO at a reported valuation of roughly $50 billion, exceeding Baidu’s own market capitalization. Prospective investors have reportedly been asked to commit to future chip purchases worth three to seven times the value of their equity subscriptions.

Model developers are moving closer to the hardware as well. Huawei’s role has extended well beyond supplying processors. Huawei executive Eric Xu said at a September 2025 keynote that Huawei Cloud engineers had “worked around the clock” to support DeepSeek’s expanding traffic, with research teams from both companies collaborating between January and April.

Reuters later reported that DeepSeek gave Chinese chipmakers early access to V4 ahead of its April 2026 release, allowing them to optimize software and kernels before launch, while Nvidia and AMD did not receive the same access. Hardware support had shifted from post-launch compatibility toward coordinated development.

The business case is becoming easier to see. The Information reported that daily token usage for Z.ai’s recently released GLM-5.2 increased 27-fold during its first week on Vercel’s model platform. As model usage grows, inference shifts from a technical problem to a recurring operating expense, making even small reductions in cost per token increasingly valuable.

Export controls explain why the market opened. They do not explain why nearly everyone is now trying to enter it. The same pattern has emerged in the United States, where Amazon designs Inferentia, Google builds TPUs, Microsoft develops Maia and Meta deploys MTIA despite having unrestricted access to Nvidia hardware. Whatever is happening in China, export controls cannot be the whole story.

The Global Inference Experiment

If China’s inference boom raises one obvious question, it is why every AI company doesn’t simply build its own chip.

Designing silicon turns out to be only half the problem. The harder challenge is keeping it busy.

Training optimizes model parameters through repeated forward and backward passes over large datasets. The workload changes materially from one model generation to the next, making flexibility valuable. Inference executes a trained model repeatedly after deployment. Prompt processing, or prefill, consists largely of highly parallel matrix operations, while autoregressive decoding generates one token at a time and is often limited less by compute than by moving model weights and stored context through memory.

That changes what matters. Peak floating-point operations per second, or FLOPS, reveal surprisingly little about the economics of serving AI. Memory capacity and bandwidth, context length, batching, latency, networking, compiler quality, utilization and power consumption all shape the cost of each token. Nvidia increased memory bandwidth from 3.35 terabytes per second on the H100 to 4.8 on the H200 and 8 on the B200 because keeping processors supplied with data matters as much as adding more arithmetic units. A processor waiting for model weights is not productive inference capacity.

Nvidia reveals Blackwell B200 GPU, the 'world's most powerful chip' for AI  | The Verge
Nvidia B200 chip.

DeepSeek illustrates the point. The company says V4 requires 73% fewer inference FLOPs per token and 90% less KV-cache memory than V3.2, attributing the improvement primarily to algorithmic changes rather than new hardware. It has separately said a further significant reduction in V4 Pro pricing will depend on Huawei’s Ascend 950 supernodes reaching scale in the second half of 2026. Software optimization and hardware deployment are reducing costs at different stages of the serving stack.

Nvidia’s biggest advantage is therefore not its hardware alone.

A cloud provider can use the same GPU fleet for training, post-training, batch inference and real-time serving, shifting capacity as models evolve and customer demand changes. CUDA lowers migration costs, TensorRT-LLM optimizes execution, Dynamo schedules prefill and decode, while NVLink and Nvidia’s networking tie processors together at rack scale. Because the same platform is available through every major cloud provider, expensive hardware is easier to keep fully utilized.

That helps explain why Nvidia’s largest customers are also building their own chips.

Amazon knows what runs inside AWS, allowing it to develop Inferentia for inference while newer Trainium products span both training and serving. Google built TPUs first for its own services before making them available through Google Cloud, with Ironwood designed specifically for inference. Microsoft designed Maia 200 around token generation, equipping it with 216GB of HBM3e memory, 7 terabytes per second of memory bandwidth, FP4 and FP8 support, and a 750-watt thermal envelope. Microsoft says Maia delivers 30% better performance per dollar than the latest hardware elsewhere in its fleet.

Meta has gone even further. It has deployed hundreds of thousands of MTIA accelerators across recommendations, advertising and consumer applications while making its roadmap with Broadcom increasingly inference-first. It does not need to convince outside customers to adopt a new architecture because the workloads already exist inside its own products. The investment is recovered through lower infrastructure costs and higher engagement rather than chip sales.

Groq illustrates the alternative. Its low-latency architecture proved there was room for specialized inference hardware, but its technology ultimately found a larger market inside Nvidia’s ecosystem than as a stand-alone platform after Nvidia licensed it on a non-exclusive basis and incorporated the Groq 3 LPX accelerator into Vera Rubin.

The dividing line is turning out to be surprisingly simple. Companies that control large, predictable workloads increasingly build specialized hardware. Companies that do not continue buying flexibility from Nvidia.

Custom AI chips are not primarily a semiconductor problem. They are a utilization problem.

Why the Incentives Converge

Once inference becomes a utilization problem, China starts to look different.

As we’ve argued since export controls first took effect, they created the opening—but not the market. In April 2025, the United States imposed an export-license requirement on Nvidia’s H20, the last China-compliant datacenter GPU then available at scale. Existing Nvidia systems remained installed and some supply continued, but Chinese companies could no longer plan future capacity around unrestricted access to the incumbent.

That does not mean China is abandoning Nvidia altogether. The Information reported in July that Beijing was considering allowing selected AI companies to import fewer than 200,000 H200 processors, less than half the quantity requested. The chips would reportedly be reserved for training while domestic processors handled inference. China has not announced a final policy, and Nvidia has said any licensed H200 sales into the country would be limited. Whether or not the proposal is adopted, the direction is clear: preserve access to frontier hardware where it remains essential while pushing recurring inference workloads toward domestic systems.

The market is already moving that way. IDC estimates China shipped roughly four million AI accelerators in 2025. Nvidia’s share fell to about 55% from roughly 95% two years earlier, while domestic suppliers reached 41%. TrendForce expects Nvidia’s share to fall again in 2026, with Huawei- and Cambricon-led vendors accounting for more than half the market and chips designed by Chinese internet platforms supplying much of the remainder. Those figures cover accelerators broadly rather than inference deployments alone, but they illustrate how quickly domestic hardware has expanded.

China's semiconductor industry is racing to catch the West's
Bernstein estimates of China AI chip market share breakdown in 2026.

Demand is moving just as quickly. As we’ve written about before, China’s National Data Administration said average daily token consumption exceeded 140 trillion in March, more than 1,000 times the level reported in early 2024 and 40% higher than at the end of 2025. ByteDance alone reported processing 180 trillion tokens a day in June, roughly ten times the volume a year earlier.

That demand increasingly has somewhere to go. ByteDance has discussed purchasing at least 50,000 inference chips from Iluvatar CoreX and has increased the share of its 2026 AI-chip budget allocated to domestic silicon. Alibaba can steer workloads toward T-Head. Baidu can do the same with Kunlunxin. Once a company controls both the workload and the deployment platform, specialized hardware becomes much easier to justify.

As we’ve also written about before, pricing reinforces the same trend. DeepSeek launched V4 Pro at $1.74 per million input tokens and $3.48 per million output tokens before making a 75% promotional discount permanent, reducing cache-miss pricing to $0.435 and $0.87. Kimi, Qwen and GLM flagship models cluster around $0.50 to $1.50 per million input tokens and $3 to $5 per million output tokens. On one benchmark compiled by our partner Weijin Research, DeepSeek’s inference costs were roughly one-twelfth those of the leading U.S. alternative. At those prices, even modest improvements in serving efficiency translate directly into lower prices, higher margins or broader deployment across existing products.

None of these forces is unique. Export controls, custom chips, hyperscaler demand and aggressive pricing all exist elsewhere. What makes China unusual is that they reinforce one another. Policy, economics and industrial strategy all point toward the same outcome: more domestic inference, more specialization and more companies willing to move down the stack.

Engineering the Stack

The term “model-chip co-design” has become shorthand for a wide range of engineering work. In practice, it encompasses everything from adapting a model to an existing processor to designing the processor itself. Those activities are often discussed together, even though they require different capabilities, different investments and produce very different economic returns.

The different types of co-design. Tech Buzz China.

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Tech Buzz China · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture