PodBrowser
a16z

Fei Fei Li: The Race to Build World Models For AI

Friday, 4 September 2026 · 3 min read · Listen to the episode ↗

In this episode, Fei-Fei Li discusses Atlas, a revolutionary model for spatial intelligence that enhances world modeling through advanced view prediction techniques. She highlights Atlas's ability to generate and reconstruct environments with minimal input frames, marking a significant leap in understanding geometry and spatial structure. The conversation also touches on the challenges of scaling the model, the importance of simulation in robotics, and the need for effective training policies to prepare robots for real-world applications.

Fei-Fei Li introduces Atlas, a groundbreaking approach to spatial intelligence that enhances world modeling through new view prediction. Justin Johnson elaborates on Atlas's capabilities, highlighting its ability to generate, reconstruct, and simulate environments using camera-conditioned generation, which allows for the reconstruction of the real world from as few as one to 100 frames. This model can create bullet time videos and perform robotic simulations, marking a significant advancement by combining generation and reconstruction in a single framework.

Li emphasizes the importance of spatially grounded meaning for each frame input, which enables precise reconstruction and represents a major leap in understanding geometry and spatial structure. She argues that unifying pixel generation and reconstruction is essential for spatial intelligence, which must facilitate generation, reasoning, editing, and interaction within spatial contexts. In contrast to the previous model, Marble, which had limitations in output representation, Atlas offers a more integrated approach to output modalities.

The design process for Atlas involved an early decision to bifurcate modalities, and scaling the model required numerous iterations and smaller experiments to build confidence. Li notes that the model can reduce dense captures from hundreds of images to as few as three, showcasing the interplay between generation and reconstruction in addressing 3D reconstruction challenges. She also points out that even expert captures can overlook details, underscoring the necessity for generative capabilities in models.

Li identifies training compute as the primary bottleneck for scaling Atlas, expressing confidence in the scaling law hypothesis while acknowledging uncertainty regarding the speed of achieving success. The architectural approach is still in its early stages of scaling potential, with critical choices in architecture and data mixtures significantly influencing outcomes. Performance improves markedly with larger model sizes and extended training durations, although the model's release was constrained by a deadline rather than limitations in scale or data.

Ben discusses World Labs' historical focus on ensuring consistency in 2D and 3D images for creatives, noting that generative view synthesis is a relatively new challenge in AI. Users are increasingly interested in building collections of assets and modeling worlds, moving away from reliance on transient generated content. He highlights the labor-intensive nature of translating feedback into 3D representations. Li identifies the acquisition of Cinex as pivotal for connecting real and simulated environments in robotics, stressing that data collection remains the most significant challenge in the field, with environment randomization being crucial for effective robotic training.

Li underscores the critical role of simulation in robotics, asserting that it is essential for training robotic policies. She emphasizes the need for exposure to potential failures during deployment to ensure effective training outcomes. Li outlines two approaches to simulation: classical simulation and data-driven simulation, explaining that learned simulators can leverage extensive data to create environments for training robotic policies.

The core thesis of world models, according to Li, is that a model should be capable of generating and simulating worlds. This capability is vital for understanding how the world reacts to actions and for envisioning appropriate actions to take. Li mentions that the launch of a new model this year received overwhelmingly positive feedback, with experts praising it as fantastic and amazing. However, she acknowledges that a robust frontier foundation model for robotics has yet to be established, indicating ongoing challenges in the field.

This summary was generated from the episode transcript and can contain mistakes.