# The Cross-Pacific AI Chessboard ## The two AI ecosystems are not racing in parallel, they are wired into each other Dimension sent a subteam through Beijing, Shanghai and Hong Kong in August 2026 and came back with a picture that inverts the standard supremacy framing. American labs train the frontier. Chinese labs distil it and release the weights. American application companies then fine-tune on those Chinese weights and sell the output back to US enterprises. Capability crosses the Pacific twice before it reaches a customer. Cursor's "self-developed" Composer 2 was built on Kimi K2.5. Cognition's SWE-1.5 looks like customised GLM. Harvey post-trained a proprietary model over Kimi's open-weight K3. A year ago these companies reached for Llama. ==The substrate of American vertical AI is now Chinese open weights with American frontier intelligence latent inside it.== This makes "who is winning" close to a non-question. See [[China - Risky Bet or Unique Opportunity]] for the prior framing I had, which this materially updates. ## Scarcity is the parent of invention Export controls were designed to slow China down. What they produced instead was evolutionary pressure that created a different species of lab. Denied compute, Chinese teams routinely engineer below abstraction layers American labs can afford to ignore. DeepSeek trained V3 on 2,048 deliberately crippled H800s by programming in PTX beneath CUDA and reallocating 20 of 132 streaming multiprocessors per GPU to compensate for nerfed interconnect ([[DeepSeek Technical Differentiators]]). Dimension's read on why this is not simply a talent gap is the sharpest line in the letter, and it is an incentive argument: the marginal dollar at a US lab buys one-off compute, while the marginal systems engineer at a Chinese lab reduces the need for it. One is consumed, the other compounds. The result is a full-stack efficiency culture across kernels, optimisers, serving systems and now silicon. > [!important] The asymmetry to hold onto > Capex buys a position. Compiler and systems engineering buys a rate of change. Constraint-driven organisations accumulate the second kind. ## Open weights walk around the procurement wall Two decades of Chinese enterprise software failed to penetrate Western production stacks. The current vintage did it in eighteen months. Chinese providers went from under 2% of OpenRouter token volume to over 45%. Qwen crossed a billion Hugging Face downloads faster than any model family in history. Roughly 80% of US AI startups run at least one Chinese open-source model in production. The mechanism is simple and worth internalising: ==there is no vendor to reject when the software is free==. Procurement, vendor risk review and country-of-origin screening all assume a counterparty. Open weights have none. What China is capturing is Western *workloads*, not Western *revenue*. The inference margin flows to whoever serves the tokens ([[Inference is Eating AI Compute]]). So the fast-follow open-weight ecosystem does two things at once: it compresses the revenue layer American frontier labs depend on, and it pushes that revenue toward hyperscalers and neoclouds ([[Hyperscalers]]). Whether designed or emergent, it is [[Why Giving Away Business Models is Genius]] executed at national scale, and it is a live stress test of the moats in [[AI era Defensibility]] and [[7 Powers]]. ## Five competitors beat two The Chinese frontier is DeepSeek, Alibaba (Qwen), Moonshot (Kimi), ByteDance (Doubao) and Zhipu (GLM), close enough that rankings reshuffle quarterly. Tencent has fallen off the Pareto frontier despite its stature. Five labs competing on cadence and capability while all releasing weights is a structurally different ecosystem from two closed American labs. Western headlines focus on inter-country competition. Intra-country competition is doing at least as much of the work. This is the generic point that competitive density, not national aggregate, drives release velocity. ## Opposing constraints, and which one relaxes first The US is power-constrained and chip-rich. China is chip-constrained and power-rich. China added roughly 429 GW of power in 2024 against roughly 51 GW in the US, with well over another 400 GW coming by 2030. In the US, 241 GW of data centre expansion sits idle behind a five-year interconnect queue ([[Data Centre Energy Demand]], [[Power Infrastructure]]). On the other side, Moonshot suspended new K3 signups within two days of launch for lack of inference capacity. > [!note] The variable that decides the middle term > US generation and interconnection versus SMIC and Huawei's yield and packaging curves. Everything else in the next three years is downstream of which curve bends faster. ## Revenue is still two orders of magnitude apart Anthropic's run rate passed $65B at the end of July, OpenAI's hit $40B in August. The strongest Chinese model revenue is ByteDance's video models at roughly $2-3B annualised, with the pure-play labs in the hundreds of millions and DeepSeek approaching $500M. Doubao has 345M monthly users, more than Qwen and DeepSeek combined, and generates under RMB 1M per day, essentially all commerce commissions. The interesting part is the pragmatism gap. Chinese labs will monetise through commerce and ads, which American labs have been reluctant to do out of something closer to aesthetic preference than economics. Distribution at 345M users eventually finds monetisation, as Chinese consumer internet did a decade ago. Valuations run the opposite way. Moonshot closed at $35B on roughly $300M ARR, about 115x, and is circling $50B pre-money ahead of a Hong Kong IPO, against roughly 20x for Anthropic. Five to ten times richer on revenue one to two orders of magnitude smaller. Zhipu and MiniMax have round-tripped violently in public markets since listing. This reads as a less mature public market pricing frontier software with consumer-software instincts, and the same dislocation visible in China and Hong Kong biotech indices. ## The verification fork The letter's most useful strategic question is what the binding constraint on self-improvement turns out to be. If self-improvement scales with compute, the US position is strong: two labs each running about 2 GW today, more than 5 GW apiece under contract, roughly 30 GW targeted by 2030. China will not have the silicon to run compute-bound autocatalytic loops ([[recursive self-improvement]]). If the constraint is instead ==verification quality==, or if timelines simply extend past the point where Chinese low-level optimisation and domestic silicon reach adequacy, the last eighteen months continue and accelerate: US frontier capability keeps commoditising, Chinese open weights keep spreading, and margin keeps migrating to inference providers. Dimension flags a dispassionate comparison of US and Chinese evaluation infrastructure as possibly the highest-return analysis left to do, and I agree. This is exactly the ground [[AI Verification]], [[Evals]] and [[Domain Experts as Eval Builders]] sit on, and it is under-owned. The data layer supports the verification read. The Chinese upstarts Dimension met were building evaluation and verification infrastructure rather than annotation services. UniPat, founded December 2025 by a Peking University PhD student, publishes benchmarks American labs now cite in their own model reports and claims strong results from a small model trained on its own rubric-based supervision. ==Substituting supervision quality for compute is the same adaptive response as writing PTX by hand.== Several of these teams are under 24 months old and already at $100M+ revenues. ## The API shutter scenario There is a plausible world where US frontier labs close API access to state-of-the-art models. The existing status quo is good enough for most programmatic enterprise use, so future high-end models get priced on value rather than usage. The labs trade high-volume API revenue for lower-volume, higher-priced access, and in return cut off the distillation channel that feeds the commoditisation loop ([[Distillation]]). That would force genuinely American-scale pretraining inside China and put real pressure on Eastern compute chokepoints. Dimension believes this is already under way for next-generation capability in science, engineering and mathematics. ## Why blunt decoupling fails American and Chinese technologists are already speaking to each other through distillation, open weights, data labelling, inference, and increasingly through software harnesses and agentic search infrastructure. American applications run on Chinese weights. Chinese labs train on rented Nvidia capacity in Malaysia, Thailand and Japan, distilled on American model outputs. Expert data flows both ways. Entity listings, hosting bans and procurement prohibitions land on the American side of this loop first. They stifle American innovation and cede category leadership, which then invites the fast-follow Chinese competitor the restriction was meant to prevent. And released weights cannot be recalled, so any decoupling applies only to future flows. For my own positioning work this cuts in one direction: sovereignty arguments built on nation-state separation are weaker than they look, and the defensible version is about jurisdiction, control planes and chokepoints rather than provenance ([[Sovereign AI Positioning]], [[The Resolver as Sovereignty Chokepoint]]). ## What I take from this The cheap conclusion is that China is catching up. The better one is that the stack has become interdependent enough that supremacy language no longer describes anything measurable, unless tractable recursive self-improvement arrives and changes the structure of the board. Three things I would hold as live positions: Constraint-driven engineering cultures compound in a way capex does not, which argues for weighting systems and compiler depth heavily when assessing any lab or infrastructure company. Evaluation and verification infrastructure is the least-analysed high-consequence layer on either side of the Pacific, and the Chinese cohort building it is younger and scaling faster than the Western equivalent. Power and silicon are the two clocks. Everything else in the middle term is a derivative of which one relaxes first.