In a significant leap for artificial intelligence, Fei-Fei Li’s startup, World Labs, has announced the release of Atlas, the world’s first multimodal world model. This groundbreaking AI system promises unprecedented control over image and video generation, capable of producing content with pixel-level camera accuracy and reconstructing scenes in 3D.
Pixel-Perfect Control and 3D Reconstruction
Developed from the ground up, Atlas is designed to process multimodal inputs, including complex camera movements, and translate them into intricate 3D views. By precisely positioning multiple input views within its spatial context, Atlas empowers users to generate images and video frames with exact camera control. This capability enables stunning visual effects, including the highly sought-after ‘bullet time’ cinematic style.
Beyond generating novel views, Atlas excels at 3D reconstruction. It can output explicit 3D models from as few as one input image, surpassing the performance of leading open-source reconstruction models. Users can specify any camera position and angle, and Atlas will render reference images that seamlessly match the content and geometry of the input. This allows for smooth expansion of images and autonomous imagination of unseen parts of a scene.
A Versatile AI for Diverse Applications
Atlas is engineered to handle a wide spectrum of tasks spanning world generation, reconstruction, and simulation:
- Camera Control Generation: Atlas produces images and videos with pixel-precise camera control. It can output videos up to one minute long in 1440p resolution, offering remarkable flexibility for creative professionals.
- Spatial Reconstruction: The model can reconstruct realistic scenes from a single image to dozens of inputs. It not only generates frames from new perspectives but also provides explicit 3D outputs, outperforming specialized 3D reconstruction models.
- Spatiotemporal Simulation: By modeling both space and time from input videos, Atlas can re-compose scenes for enhanced dramatic effect. It also supports realistic-to-simulation workflows for robotics applications.
- Image Generation: Atlas supports text-to-image generation and the creation of 360-degree panoramas. It adeptly handles complex prompts, renders text accurately, and can produce visuals in a variety of artistic styles.
Early Access and Future Rollout
World Labs has indicated that select partners have already received early access to Atlas. The company plans to open up broader early access in the coming weeks, signaling an exciting period of adoption and integration for this advanced AI technology.
The release of Atlas by World Labs, spearheaded by AI pioneer Fei-Fei Li, marks a pivotal moment in the development of AI. Its ability to understand and generate complex visual information with such precision opens up new frontiers for creativity, research, and industry applications.









