What do we know about Atlas, the AI model created by Fei-Fei Li’s World Lab?
The aim is to construct a coherent representation of space, reconstruct three-dimensional environments and simulate their evolution over time.
Having taught machines to recognise what they see, Fei-Fei Li now wants to teach them to imagine what they would find if they changed their perspective. Atlas, the new model unveiled by World Labs in early September, was developed with this ambition in mind: to build a coherent representation of space, reconstruct three-dimensional environments and simulate their evolution over time.
This is a different objective from that of chatbots. A large language model learns the statistical relationships between words and generates the most likely continuation of a text. Atlas attempts to do something similar with space: it combines text, images, video sequences, camera positions and depth information to predict what should be visible when viewing a scene from a different angle.
World Labs describes it as an ‘omni-world model’ for spatial intelligence. Technically, it is a multimodal autoregressive diffusion transformer: an architecture that combines elements of language models and image generators. All inputs are placed within a shared spatial context, and the system progressively generates new views, whilst striving to maintain consistency in geometry, objects and perspective. World Labs
Who is Fei-Fei Li
The name behind World Labs explains why Atlas is attracting so much attention. Fei-Fei Li is a professor of computer science at Stanford, co-founder of the Stanford Institute for Human-Centred Artificial Intelligence and former director of the Stanford AI Lab. Between 2017 and 2018, she also served as vice-president at Google and chief scientist for AI and machine learning at Google Cloud. Stanford HAII Her best-known contribution is ImageNet, the vast repository of annotated images that has made it possible to measure and accelerate progress in computer vision. It was a competition based on ImageNet that, in 2012, highlighted the quantum leap achieved by deep neural networks.
What can Atlas do? From a photograph to a three-dimensional environment
As stated on the Atlas website, it can take one or more photographs as its starting point and generate what a camera would see as it moves around the scene. The difference compared to many video generators is that the camera’s path is not merely described in words: position, direction and movement are native inputs to the model. According to World Labs, this allows for precise control over framing and trajectories and enables the production of videos up to one minute long, with a maximum resolution of 1440p.



