Science

What do we know about Atlas, the AI model created by Fei-Fei Li’s World Lab?

The aim is to construct a coherent representation of space, reconstruct three-dimensional environments and simulate their evolution over time.

4' min read

Translated by AI
Versione italiana

4' min read

Translated by AI
Versione italiana

Having taught machines to recognise what they see, Fei-Fei Li now wants to teach them to imagine what they would find if they changed their perspective. Atlas, the new model unveiled by World Labs in early September, was developed with this ambition in mind: to build a coherent representation of space, reconstruct three-dimensional environments and simulate their evolution over time.

This is a different objective from that of chatbots. A large language model learns the statistical relationships between words and generates the most likely continuation of a text. Atlas attempts to do something similar with space: it combines text, images, video sequences, camera positions and depth information to predict what should be visible when viewing a scene from a different angle.

Loading...

World Labs describes it as an ‘omni-world model’ for spatial intelligence. Technically, it is a multimodal autoregressive diffusion transformer: an architecture that combines elements of language models and image generators. All inputs are placed within a shared spatial context, and the system progressively generates new views, whilst striving to maintain consistency in geometry, objects and perspective. World Labs

Who is Fei-Fei Li

The name behind World Labs explains why Atlas is attracting so much attention. Fei-Fei Li is a professor of computer science at Stanford, co-founder of the Stanford Institute for Human-Centred Artificial Intelligence and former director of the Stanford AI Lab. Between 2017 and 2018, she also served as vice-president at Google and chief scientist for AI and machine learning at Google Cloud. Stanford HAII Her best-known contribution is ImageNet, the vast repository of annotated images that has made it possible to measure and accelerate progress in computer vision. It was a competition based on ImageNet that, in 2012, highlighted the quantum leap achieved by deep neural networks.

What can Atlas do? From a photograph to a three-dimensional environment

As stated on the Atlas website, it can take one or more photographs as its starting point and generate what a camera would see as it moves around the scene. The difference compared to many video generators is that the camera’s path is not merely described in words: position, direction and movement are native inputs to the model. According to World Labs, this allows for precise control over framing and trajectories and enables the production of videos up to one minute long, with a maximum resolution of 1440p.

The model can also combine different photographs, position them virtually within a space and imagine corridors, doors or passageways connecting them. It is a powerful feature, but also the point at which reconstruction and generation risk becoming confused: when information is lacking, Atlas does not uncover the hidden reality, but rather constructs a plausible solution.

The company states this quite openly. With just a single image, the system has to work out much of what it cannot see; by adding photographs, the reconstruction becomes more accurate. In some cases, two or three views are sufficient, whilst to reproduce a location with greater precision, Atlas can process more than a hundred.

The result can be output as a video, a point cloud or using 3D Gaussian splats – a technique that represents a scene using millions of coloured three-dimensional elements and allows it to be explored in real time. The most obvious applications are in cinema, video games, design, special effects and virtual environments.

The Challenges Facing World Models

World Labs claims that Atlas outperforms certain video models in its ability to track complex camera movements, as well as various open-source systems specialising in 3D reconstruction. The comparisons were carried out, respectively, using human evaluators and academic benchmarks based on reconstruction error.

However, these are results presented by the company itself. No full scientific paper has been published setting out all the details of the model, the training data and the evaluation procedures, nor is Atlas yet freely available: it is entering an early access phase with a select group of partners.

Loading...

Who works on the World Models?

Atlas is not alone. Google DeepMind has developed Genie 3, which is capable of creating navigable environments in real time at 720p and 24 frames per second from a single prompt, whilst maintaining them consistently for several minutes: virtual worlds in which to train and evaluate autonomous agents. Microsoft Research, with Muse — based on the World and Human Action Model (WHAM) — generates gameplay sequences by predicting how the scenario will change in response to the player’s commands; for now, it is primarily aimed at video game design, rather than robotics. Nvidia Cosmos, on the other hand, offers models and tools specifically designed for ‘physical AI’, with which to produce synthetic videos and simulations for robots and autonomous vehicles. Meta’s V-JEPA models are working in the same direction, learning a representation of the world by predicting how a visual scene will evolve. The difference lies in the starting point: Genie focuses on generated interactive worlds, Muse on the dynamics of video games, Cosmos on data for autonomous machines; Atlas focuses primarily on the controllable spatial reconstruction of the real world.

Copyright reserved ©
  • Luca Tremolada

    Luca TremoladaGiornalista

    Luogo: Milano via Monte Rosa 91

    Lingue parlate: Inglese, Francese

    Argomenti: Tecnologia, scienza, finanza, startup, dati

    Premi: Premio Gabriele Lanfredini sull’informazione; Premio giornalistico State Street, categoria "Innovation"; DStars 2019, categoria journalism

Loading...

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti