Executive Summary
Four unrelated sources converge on one substrate question: who controls persistent, general-purpose compute, and how honestly is its capability disclosed to the party assuming the risk. Meta ships every consumer a persistent Linux VM behind a mascot. China's hyperscalers double capex into a 24GW+ base that Western public filings undercount by up to 15x. Runway and Nvidia argue that video models trained at scale become world models, converting robotics data acquisition from a physical-logistics cost into a GPU line item. The common mechanic is capability arriving faster than the disclosure, accounting, or policy layer around it. For enterprises, the near-term exposure is governance (BYO-agent with root-level reach), and the near-term planning error is sizing Chinese AI capability or physical-AI cost curves from the wrong denominator.
What Changed
SemiAnalysis's new bottom-up China Datacenter Model (1,000+ facilities, 60+ operators) is the hard data point: 24GW+ live, larger than EMEA or APAC ex-China, with ~50GW dated or announced. The reason the market missed it is structural, not analytical laziness. GDS and VNET, the only US-listed proxies, capture about a third of ByteDance and Alibaba orders, and ByteDance, roughly a fifth of delivered capacity and almost entirely leased, files nothing. 2Q26 BAT capex hit $20B, more than double YoY, with all three at negative free cash flow simultaneously for the first time. That is not optionality spending. It is AI as the only growth lever left while WeChat MAU grows 2%, Baidu revenue falls 3%, and Alibaba commerce falls 7%.
Second change: world models stopped being research framing. Runway found that video pretrained on third-person footage transfers to robotic manipulation with hundreds of hours of embodiment-specific fine-tuning rather than hundreds of thousands, with measurable sim-to-real correlation and no hand-built 3D simulator. Nvidia's kitchen-manipulation argument is the same economics from the seller's side.
Cross-Expert Synthesis
Runway and Nvidia agree on the mechanism and disagree implicitly on the moat. Nvidia's version keeps simulation fidelity as the scarce asset, which sustains compute demand. Runway's version says the scarce asset is abundant third-person video plus scale, which makes the robotics simulator a byproduct of a media business. If Runway is right, robotics platform advantage sits with whoever has the largest video pretraining corpus, not whoever has the best physics engine or the most field-collected teleoperation hours. Both agree the cost curve moves to compute, which is predictable, from rigs and field testing, which are not.
Willison and SemiAnalysis sit on opposite ends of the same visibility failure. Willison's is downward: consumers cannot see what their agent can do because the interface deliberately understates it. SemiAnalysis's is outward: strategists cannot see what China has built because the tenants and landlords are private. Both punish anyone who models from the artifact they can see rather than the substrate underneath.
The sharpest cross-source tension is sovereignty. Runway flags that most top-ranked open video models are Chinese, motivating the Nvidia-backed Cosmos Coalition. SemiAnalysis explains why: China is chip-gated, not power- or permit-gated, delivering 100MW in ~12 months via prefab (T-Block, CUBE 5.0) against 24-30+ months in the US. Constrain the input, and the output concentrates in video and open weights, exactly where compute per unit of capability is cheapest.
Enterprise Implications
Agentic tools with persistent execution environments are now a consumer category, which means they are already a shadow-IT category. The exposure is not a novel exploit; it is an opaque long-lived process holding credentials with desktop reach, governed by UX copy. Expect enterprise SaaS vendors to copy Meta's disclosure posture, not correct it.
Any robotics or physical-automation vendor evaluation that weights hardware specs over sim-to-real transfer maturity is scoring the wrong variable. Program cost will track GPU pricing.
Runway's capability-absorption pattern, prompt rewriting and multi-shot orchestration moving end-to-end into base models, prices the durability of harness and orchestration tooling built on today's video APIs at roughly two years.
What To Do About It
- Write acceptable-use policy for agents with persistent execution environments now, scoped to credential handling, data residency, and desktop reach. Policy written after the first incident will be written by legal, not by you.
- Inventory which employees already run agentic tools with shell access on machines touching corporate data. Treat this as endpoint discovery, not a survey.
- Re-baseline any China AI capability assumption that traces back to GDS/VNET filings or public capex disclosures. Add the ~4GW-by-2029 offshore leasing trend and Chinese GPU rental from Western clouds, which direct capacity figures exclude.
- In robotics RFPs, require a demonstrated sim-to-real correlation metric on a manipulation benchmark. Vendors without one are still paying for physical iteration.
- Do not build durable tooling around video-model orchestration harnesses. Assume absorption.
Customer question worth asking directly: "For every AI feature your vendors shipped this year, can you state the permission scope and process lifetime?" Most CIOs cannot, and the gap is the engagement.
Risks and Blind Spots
Nvidia and Runway both benefit commercially from the world-model thesis; the sim-to-real claim is vendor-reported and not independently benchmarked. Interface-as-video-model is not cost-competitive with HTML rendering and may never be outside narrow high-value cases. SemiAnalysis's 24GW is live capacity, not utilized capacity or delivered FLOPs, and the ~50GW announced figure is the softest number in the set. China's capex surge is defensive, funded by negative free cash flow against decaying core businesses, which is a fragility as much as a signal.