World Labs Inc., the high-profile and well-funded artificial intelligence startup co-founded by the renowned computer vision pioneer Fei-Fei Li, has just dropped Atlas, which promises to be a game-changer in the world of “world models.”
In a blog post, World Labs explained that Atlas is a breakthrough multimodal world model that aims to bridge the gap between simulated environments and physical reasoning. The company said Atlas is an “omni model,” which excels in creating expansive and highly-detailed simulated 3D environments from a single image input, with precise camera control, enabling them to be viewed from any angle.
The release of Atlas appears to mark the fulfillment of Li’s overarching theory of “spatial intelligence,” which is the idea that if AI is to really understand and replicate the physical world, it must be able to reason natively about everything that happens within them – including the 3D objects and people, the environments, how these elements interact and the spatial consequences of those interactions.
Introducing Atlas:
The world’s first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time. pic.twitter.com/o0qeGubi19
— World Labs (@theworldlabs) September 1, 2026
Li, who previously served as the director of Stanford University’s AI Lab and co-founded the Stanford Institute for Human-Centered AI, launched World Labs in February 2024. At the time, she said that it’s impossible to create “artificial general intelligence,” or an AI system that surpasses the cognitive capabilities of humans, if they don’t have spatial intelligence.
Therefore, World Labs set out to develop physics-aware world models capable of understanding, reasoning with, and taking actions within the physical world. It’s an idea that has won a lot of favor with investors, with the startup raising $1.2 billion in funding from backers including Nvidia Corp., Advanced Micro Devices Inc. and Autodesk Inc.
Atlas is built on what World Labs says is a multimodal autoregressive diffusion transformer architecture that’s far more complex than early video generators such as OpenAI Group PBC’s Sora AI. Whereas traditional video generators rely on imprecise text prompts to control camera angles and movements, Atlas ingests camera trajectories and geometry as native inputs, enabling extremely precise control in terms of perspective.
Early adopters have already been posting the results of their experiments with Atlas on platforms like X, and the first impressions look extremely positive:
This scene was generated from one input image. Atlas fills the remaining gaps.
→ atlas (world labs) → spark.js → three.js pic.twitter.com/htTVjg0iP4
— Ian Curtis (@XRarchitect) September 1, 2026
What’s really impressive is that Atlas can create such incredibly realistic simulations from just a single 2D image. From that input, it can generate up to a minute of 1440p video that maintains rigid geometric consistency while being viewable from any angle. It can also output 3D assets such as point clouds and 3D Gaussian splats, enabling it to combine video, text, camera poses and depth maps and create shared spatial context.
World Labs showed off a number of stunning examples, and said Atlas has applications in creative visual effects and game design. But the startup said its real target is creating extremely accurate simulations for robotics training. Atlas supports “c” workflows, where developers can use an ordinary smartphone camera to capture a physical space, and then reconstruct that scene as a 3D simulation that robots can navigate through and interact with.
As the robot moves around and manipulates objects within the space, Atlas will generate highly precise RGB images and depth sensor readings. Add to that its 360-degree camera view, and it creates the perfect environment for robotics developers to safely test their hardware in almost any kind of scenario, with different object arrangements, lighting conditions and so on.
World Labs said it has tested Atlas on a number of benchmarks against some of the industry’s top video generation and 3D reconstruction models. In a blind human evaluation focused on assessing camera-path adherence, Atlas was overwhelmingly preferred compared to competitors such as Gemini Omni Flash and FLUX. In addition, Atlas outperformed a number of top open-source models in terms of its ability to reconstruct 3D geometry from sparse inputs.
These results help to distinguish Atlas in an incredibly competitive market for world models. World Labs faces some strong competition, with the likes of Odyssey developing interactive world simulations, Yann LeCun’s AMI Labs focused on physical planning architectures and Niantic Spatial building advanced geospatial mapping systems. On the other hand, Atlas is a single base model that tries to unify all of these capabilities.
As impressive as its demonstrations are, the real test for Atlas will come with wider availability, when users get the chance to replicate World Labs’ initial results in complex, real-world scenarios. However, though World Labs says Atlas is available in early access now for select enterprises, it has not yet said when it will be generally released.





