On August 28, Tencent published the preview of Hy4, a new mixture-of-experts model with a 770-billion-parameter backbone, 49 billion active parameters per token and a one-million-token context window. One week earlier, DeepSeek had added multimodal agent capabilities to its V4 family. Those releases followed Moonshot AI’s 2.8-trillion-parameter Kimi K3 in July and Alibaba’s Qwen3.8 family in August.
The cadence matters more than any single benchmark result. China’s AI story is no longer a DeepSeek-style cost shock produced by one unusually efficient model. It is becoming a broad model portfolio spanning frontier-scale systems, lower-cost inference, open weights, multimodal reasoning, coding and research agents, and increasingly mature cloud distribution.
That strengthens the case for a more bipolar global AI industry. But “U.S. AI versus China AI” is too simple a description of what is forming. The model layer remains surprisingly portable. The deeper separation is appearing underneath it—in compute, cloud infrastructure, capital, policy and deployment sovereignty.
Global AI is moving toward two major technology stacks, but not two sealed ecosystems.
China’s model race is no longer a one-model story
DeepSeek changed perceptions of Chinese AI by demonstrating that lower cost could itself be a strategic advantage. The releases that followed have changed the question. Chinese developers are now pushing both the lower and upper ends of the market.
Kimi K3 illustrates the change at the high end. Moonshot describes the model as the first open model in the three-trillion-parameter class. It combines 2.8 trillion total parameters, native vision and a one-million-token context window, with an architecture designed for long-horizon coding, knowledge work and reasoning. Moonshot itself acknowledges that K3 still trails the strongest proprietary systems in overall user experience. That caveat is useful: China’s progress does not require claiming universal frontier leadership.
Alibaba is following a similar direction. Qwen3.8-Max pushed the Qwen family further into coding, professional work and long-horizon agent tasks, while Alibaba subsequently released the closely related Qwen3.8-2.4T-A95B weights. Qwen is no longer just a collection of smaller open models designed to maximize distribution. It now spans lightweight deployment through to models operating near the upper end of the capability curve.
DeepSeek has moved in parallel. V4 Flash and V4 Pro increasingly emphasize agentic execution rather than conventional chat. The August V4 Pro update added stronger tool-using capabilities and native support for the OpenAI Responses API format, while an experimental V4 Flash Vision release extended the family into multimodal agents.
Tencent’s Hy4 preview adds another credible participant. Its importance is not that Tencent has necessarily produced the best Chinese model. It is that the number of Chinese organizations capable of releasing large, competitive models continues to increase. Alibaba, DeepSeek, Moonshot, Tencent, Z.ai and MiniMax are creating something that did not exist during the first phase of generative AI: depth behind the Chinese frontier.
This makes the competitive mechanism different from the one implied by a simple benchmark race. A country with one strong model has a product. A country with several model developers, hyperscale cloud providers, domestic accelerators, developer tools and recurring capital investment begins to have an ecosystem.
Open weights change the economics before they change geopolitics
Chinese AI also differs from much of the first generation of U.S. frontier AI in how aggressively models are distributed. Kimi, Qwen, DeepSeek, Tencent and several other Chinese developers have made open-weight releases an important part of their strategies.
That matters economically because model weights can separate the model developer from the inference provider. An enterprise does not necessarily need to buy every token from the organization that trained the model. It can self-host, use a third-party inference provider, optimize the serving stack or deploy the same weights across several environments.
OpenRouter provides a useful illustration. In July, the platform reported that DeepSeek V4 Pro was available through 16 different inference providers. The same underlying model showed material differences in price, throughput and availability depending on who served it. Once model weights become portable, part of the competitive advantage migrates from intellectual-property access toward inference engineering, accelerator utilization, networking, memory, software optimization and power cost.
This does not make open weights a uniquely Chinese strategy. NVIDIA is expanding its Nemotron open-model family and explicitly positions open models as a substrate for enterprise and agentic workloads. U.S. developers can therefore respond to China’s distribution model without abandoning the broader U.S. compute ecosystem.
The result is an unusual competitive structure. National separation can increase at the infrastructure layer even while technical interoperability increases at the model layer.
Model adoption is starting to look like an infrastructure cycle
Competitive models matter only if they generate usage. Here the evidence is becoming more consequential.
OpenRouter’s request logs show that Chinese-developed models surpassed U.S.-developed models in token volume on its platform by early June 2026. DeepSeek’s share alone roughly doubled from about 9% at the start of the year to around 18% in the January-to-June comparison. Agentic workloads accounted for much of the acceleration.
OpenRouter is not the global AI market. Its users are unusually comfortable switching models and using open-weight systems, and token volume is not the same as revenue. It should therefore be treated as a distribution signal rather than a global market-share estimate. Even with that limitation, the direction is difficult to dismiss: Chinese models are being used outside controlled benchmark environments and are competing for production workloads.
The next link in the chain is capital expenditure.
Alibaba reported that AI Cloud and Compute Services revenue reached about $7.1 billion in its latest quarter, 45% higher than a year earlier. The cloud segment’s adjusted EBITA increased 133%, while the company spent nearly $10 billion on capital expenditures during the quarter, up 75% year over year. Alibaba is therefore beginning to show both sides of the AI equation: rising infrastructure spending and a measurable revenue response.
Tencent’s second-quarter numbers are even more illustrative of the cost of competing. Capital expenditure reached RMB52.8 billion, 176% above the prior year. The company recorded negative free cash flow during the quarter as infrastructure purchases and large AI-related compute prepayments absorbed cash. Management explicitly linked the additional compute procurement to Hy model development, inference for WorkBuddy and CodeBuddy, AI functions inside Weixin and external cloud demand.
That pattern runs against the idea that more efficient Chinese models should automatically reduce global AI infrastructure demand. Lower inference cost can reduce the cost of an individual task while expanding the number of economically viable tasks. Coding agents, office agents and research systems also execute many more model calls than a conventional chatbot session.
The relevant unit is therefore shifting from cost per token toward cost per completed workflow. If lower model prices stimulate far more inference, aggregate compute demand can rise even as unit costs fall.
There is an important counterweight. Cheap tokens do not guarantee attractive model economics. If price competition moves faster than utilization and cloud monetization, model providers can create large amounts of usage while earning inadequate returns on the infrastructure required to serve it. Alibaba’s cloud growth offers evidence that monetization can work; Tencent’s cash-flow profile shows how capital-intensive the route can become.
The real bifurcation is below the model layer
The strongest case for a U.S.-China split is not found in leaderboards. It is found in the infrastructure required to keep producing competitive models.
Compute is the hardest layer to share
The United States still holds a major structural advantage in leading AI accelerators and the surrounding software ecosystem. U.S. export controls have restricted China’s access to advanced processors for years. The current regime is more nuanced than a complete embargo: since January 2026, U.S. authorities have allowed license applications for products including NVIDIA’s H200 and AMD’s MI325X to be reviewed on a case-by-case basis when specified conditions are met.
Controlled access is still very different from an unrestricted supply chain. It gives Chinese cloud companies a strong incentive to develop an infrastructure path that does not depend on future U.S. licensing decisions.
Huawei’s Atlas 950 SuperPoD shows the direction. The system displayed at the 2026 World Artificial Intelligence Conference connected 1,024 Ascend processors as a single large compute domain. Huawei’s architecture is designed to scale to as many as 8,192 NPUs, using high-bandwidth interconnects and unified memory addressing to compensate for the limitations of individual processors through system scale.
This should not be confused with proof of chip-for-chip parity. More accelerators can bring higher networking complexity, energy consumption and software requirements. China’s challenge therefore expands rather than disappears: accelerator performance, advanced memory, packaging, optical connectivity, cluster scheduling and power efficiency all become part of the competitive equation.
But the strategic objective is clear. China does not need every domestic chip to match the best U.S. accelerator individually if a sufficiently large domestic system can deliver useful training and inference capacity at an acceptable total cost.
Cloud turns model capability into an operating ecosystem
The same logic applies to cloud infrastructure. The U.S. ecosystem combines frontier model developers with AWS, Microsoft Azure, Google Cloud, specialist GPU clouds and the NVIDIA software stack. China is assembling a parallel distribution layer through Alibaba Cloud, Tencent Cloud, Huawei Cloud, ByteDance’s cloud operations and other domestic providers.
Alibaba is particularly important because its stack now reaches from semiconductor design and cloud infrastructure to Qwen models and enterprise applications. The economic logic is similar to that of U.S. hyperscalers: better models create cloud demand, cloud revenue funds infrastructure, and larger infrastructure enables the next generation of models.
A functioning feedback loop matters more than any one quarter’s benchmark ranking. The side that can repeatedly convert model quality into usage, usage into cash flow, and cash flow into additional compute has a more durable competitive position.
Policy is explicitly reinforcing the second stack
China’s 2026 government work report makes the direction unusually explicit. It calls for an expansion of the “AI Plus” initiative, faster adoption of AI agents, support for open-source AI communities, hyperscale intelligent-computing clusters, coordination between computing and electricity infrastructure, and further development of public cloud.
These policies link the model layer to infrastructure and industrial policy. Open models expand distribution. Domestic compute reduces external dependency. Public cloud lowers deployment barriers. Large-scale application creates demand that can justify further investment.
The U.S.-China competition is therefore evolving from a contest between individual laboratories into competition between capital-intensive systems.
| AI Layer | Expected Degree of U.S.-China Bifurcation | Why |
|---|---|---|
| Model weights | Medium | Open-weight releases can move across providers and jurisdictions more easily than physical infrastructure. |
| Developer APIs and tools | Low to Medium | Compatible interfaces and common agent frameworks reduce switching costs. |
| Cloud and inference | Medium to High | Data residency, procurement, local platforms and infrastructure economics encourage regional stacks. |
| Accelerators and AI servers | High | Export controls and domestic substitution increasingly separate hardware supply chains. |
| Advanced memory, packaging and interconnect | High | Physical capacity, qualification and manufacturing know-how are difficult to replicate or move quickly. |
| Policy, data and sovereign deployment | High | Security rules, data governance and industrial policy directly influence which stack can be deployed. |
Source: Sector Foundry analysis.
A bipolar AI market will not be a binary one
The most likely end state is therefore more nuanced than a technological Cold War with two closed networks.
A U.S.-led ecosystem is likely to retain advantages in leading-edge accelerator technology, global cloud distribution, proprietary frontier models and established enterprise relationships. A China-led ecosystem is increasingly capable of combining domestic models, lower-cost and open-weight distribution, local cloud infrastructure, domestic accelerators and strong policy support.
Those systems can become progressively independent without becoming technically incompatible.
DeepSeek already illustrates the point. Its models can be served by multiple providers rather than only by DeepSeek itself, and its API supports interfaces familiar to developers building around U.S.-originated standards. Qwen and Kimi weights can similarly travel farther than a Chinese data center or cloud account. NVIDIA’s own expansion of open models further blurs any attempt to equate “open” with China and “closed” with the United States.
This creates room for a third behavior outside the two core ecosystems: mixing layers. An enterprise can use a proprietary U.S. model for one high-value reasoning task, an open Chinese model for a cost-sensitive workflow, and its preferred regional cloud or private infrastructure to serve another workload. Model routers make that allocation increasingly automatic.
Regulated industries, government workloads and sensitive enterprise data will be less flexible. There, security rules, procurement policy, data localization and confidence in the underlying supply chain can force deeper alignment with one stack. Consumer software and less regulated enterprise workloads may remain considerably more fluid.
The global AI market may therefore resemble a two-pole supply structure with a partially shared software market above it.
Four tests for the two-stack thesis
- Can China keep narrowing capability gaps under a compute constraint? Model architecture and system optimization can stretch available compute, but they cannot eliminate indefinitely a large disadvantage in semiconductor manufacturing, advanced memory or energy efficiency.
- Can usage become durable revenue? Token growth matters only if cloud, API, subscription and enterprise revenue can support the capital required for the next model cycle.
- Will model portability remain politically acceptable? Security restrictions on weights, data, cloud hosting or developer access could make the software layer split much faster than it has so far.
- How aggressively will the U.S. open-model ecosystem respond? NVIDIA’s Nemotron strategy shows that open weights are not a structural advantage China can monopolize. Strong U.S. open models would narrow China’s distribution advantage while reinforcing U.S. infrastructure.
The variable worth watching is no longer whether a Chinese model can appear near the top of a leaderboard for a few weeks. It is whether Chinese providers can sustain a full economic flywheel: better models create usage; usage creates cloud and application revenue; that revenue supports compute investment; and the new infrastructure enables another generation of models.
If that loop holds despite the semiconductor constraint, a second AI stack becomes more than a geopolitical ambition. It becomes a durable industry structure.
Sources and Methodology
This analysis uses model release notes and model cards from Moonshot AI, DeepSeek, Alibaba/Qwen and Tencent; quarterly disclosures from Alibaba and Tencent; Huawei infrastructure disclosures; Chinese government policy documents; U.S. Bureau of Industry and Security export-control materials; and public model-usage data from OpenRouter. NVIDIA materials were used to assess the U.S. open-model response.
OpenRouter data is treated as a developer-platform adoption indicator rather than a measure of global AI market share. Vendor benchmark claims are treated as vendor-reported unless supported by independent evaluations. Model rankings and pricing are intentionally not used as the central thesis because both can change rapidly.
Broker research supplied for this analysis was used to identify industry questions, test competing interpretations and locate relevant public evidence. Proprietary report prose, charts and tables were not reproduced.
Research Cut-off: September 3, 2026, 20:31 KST (UTC+9)
Public primary sources
- Moonshot AI — Kimi K3 technical launch, July 2026. Kimi describes K3 as a 2.8T-parameter open model with native vision and a 1M-token context, while explicitly acknowledging that overall user experience still trails the strongest proprietary systems.
- DeepSeek API documentation — July 31, August 13 and August 21, 2026 updates. Used for V4 Flash, V4 Pro, Responses API compatibility and the multimodal agent update.
- Alibaba/Qwen — Qwen3.8, August 2026. Qwen3.8-Max and the subsequent Qwen3.8-2.4T-A95B weight release were used to assess the widening Chinese model portfolio.
- Tencent — Hy4 preview, August 2026. The official model repository describes a 770B-parameter backbone, 49B active parameters and 1M context.
- Tencent 2Q26 results, August 12, 2026. CapEx was RMB52.8 billion, up 176% YoY; the company linked compute procurement and AI-related prepayments to model, agent and cloud demand.
- Alibaba quarterly AI update, August 20, 2026. AI Cloud and Compute Services revenue reached $7.1 billion, up 45% YoY; quarterly CapEx was nearly $10 billion, up 75%.
- Huawei — Atlas 950 SuperPoD, July 2026. Used to assess China’s system-scale response to processor constraints. Huawei disclosed a 1,024-card deployment at WAIC.
- China 2026 Government Work Report. The policy agenda explicitly supports AI agents, open-source AI communities, hyperscale intelligent-computing clusters, compute-power/electricity coordination and public cloud.
- U.S. Bureau of Industry and Security, January 13, 2026. H200, MI325X and similar exports to China moved to case-by-case license review subject to conditions, supporting the article’s characterization of a controlled rather than absolute hardware separation.
- NVIDIA open-model materials, 2026. Used to test the idea that open weights are uniquely Chinese; NVIDIA continues to expand Nemotron as an open-model family for agentic workloads.
Public secondary / industry datasets
- OpenRouter, June–July 2026. More than 450 trillion tokens from January 1 through June 14 were used in its U.S.-versus-China analysis; Chinese-developed models exceeded U.S. models in token volume on OpenRouter by early June. The article explicitly treats this as platform-specific rather than global market share.
- Artificial Analysis. Used as a cross-check that recent Chinese models are competing in the frontier performance cluster rather than as the basis for a fixed ranking.

