The brief notes that no existing benchmark gives policymakers an adequate basis to approve a world model for safety-critical deployment. How big a problem is that?
Li: It’s a fundamental gap. Current benchmarks mostly measure visual quality, not whether a system understands physical dynamics. Newer benchmarks are starting to test physical reasoning—whether objects behave correctly, whether scenes stay consistent, whether skills learned in simulation transfer to real tasks.
But we’re still in a research patchwork. And leading models now achieve near-perfect scores on simulated tasks even though their real-world reliability remains limited, making those benchmarks less useful for distinguishing among systems.
Until measurement science catches up, we must continue to require rigorous field testing. We can’t shortcut real-world validation just because a system looks good in simulation.
What should policymakers do first?
Zegart: Three things. First, fund the measurement science. Direct NIST to develop evaluation methods for world models, with sector agencies defining relevant operating conditions. Without valid benchmarks, we’re flying blind on safety.
Second, use procurement strategically. Government agencies are major customers for autonomous systems, robotics, and simulation infrastructure. Every procurement should require independent testing, access to evaluate the simulation environment, and real-world validation. Make those requirements standard across federal purchasing.
Third, address the data infrastructure gap. Make datasets and public-interest simulation environments explicit priorities for the National AI Research Resource.
Li: I’d add a fourth: Invest in cross-disciplinary research and training. World models sit at the intersection of computer vision, robotics, physics simulation, and domain expertise in everything from urban planning to crisis response to national security. We need researchers and policymakers who can work across those boundaries.
Wald: Stanford HAI’s model of bringing together computer scientists, social scientists, legal scholars, policy experts, and domain specialists needs to be replicated across universities and government agencies. The governance challenges aren’t purely technical, and the technical solutions aren’t purely algorithmic. They require sustained collaboration across disciplines.
Your conclusion emphasizes that the choices made now will determine whether world models develop inside “a narrow commercial and security logic” or within a framework that serves public value. What’s at stake in that choice?
Li: We're at an inflection point: whether this technology strengthens economic resilience, expands scientific capacity, and supports safer embodied AI or whether it concentrates power, expands surveillance, and accelerates unsafe deployment in high-stakes settings.
The difference comes down to decisions being made right now: investing in shared infrastructure before concentration hardens, tying safeguards to deployment contexts, and building measurement science and public expertise.
Zegart: The window is narrow. Once proprietary systems become deeply embedded in critical infrastructure, commercial fleets, and military operations, changing course becomes so much harder.
But here’s what gives me hope: Unlike with language models, where we’re playing catch-up, we’re identifying these governance challenges while world models are still emerging from labs into early deployment. We have a chance to shape the conditions under which this technology develops before its trajectory locks in.
“The World Model and Spatial Intelligence Era: Governing AI Beyond Language” was written by the following Stanford HAI faculty and staff: Daniel Zhang, Russell Wald, Ehsan Adeli, Elena Cryst, Daniel E. Ho, Caroline Meinhardt, Jiajun Wu, Amy Zegart, and Li Fei-Fei.







