Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Thursday, 19 June 2025 · 6 min read · Listen to the episode ↗
Noam Brown discusses the scaling of test time compute in AI, emphasizing its crucial role in enhancing problem-solving capabilities within multi-agent systems. He highlights advancements in AI strategies through collaborative and competitive dynamics, drawing parallels between his work in Diplomacy and traditional game theory. The conversation also addresses the evolution of language models, including their limitations and potential, while spotlighting the implications for cryptocurrencies and blockchain technology in terms of AI's adaptability and reasoning capabilities.
Noam Brown from OpenAI shares insights from his experience with the game Diplomacy and the development of the AI bot Cicero, which ranked in the top 10% of human players. He emphasizes the importance of understanding the game for debugging and improving Cicero's strategies, which ultimately inspired his own gameplay and led him to win the World Diplomacy Championship in 2025.
In discussing human-bot interaction, Noam notes that players initially did not focus on distinguishing between bots and humans during games. He acknowledges the limitations of the language models used in Cicero, which sometimes resulted in bizarre responses, but highlights significant advancements in language models since 2022, particularly with GPT-4.0, which is capable of passing the Turing test and demonstrates improved performance due to its larger size.
Concerns about AI's persuasive abilities and the associated safety discourse are raised, to which Noam responds that the AI safety community has positively received Cicero, appreciating its controllable nature and reasoning system. He sees potential in testing various AI models in Diplomacy as a benchmark for evaluating their performance and anticipates continued rapid progress in AI capabilities.
The conversation touches on the emergence of agentic behavior in AI, particularly with the use of "oh, three," which Noam finds useful for conducting meaningful research. He addresses skepticism regarding AI's performance in unverifiable domains, citing deep research as evidence of AI's ability to excel in complex tasks. He believes that understanding the quality of outputs will enhance model performance over time.
Noam discusses the relationship between intuitive (system one) and analytical (system two) thinking in AI, suggesting that foundational capabilities must develop before higher-level functions can be effective. He notes that while smaller models like GPT-2 showed minimal results with reasoning applications, larger models demonstrated improvement. He speculates that future models, such as GPT-6, may enhance system one capabilities, potentially leading to perfect gameplay in strategic contexts.
The conversation shifts to model routing, focusing on the need for a system that can switch between fast response models and those requiring deeper thinking. The speaker notes that while a simpler model might effectively route complex problems, it could also be overconfident or misled. Looking ahead, the speaker suggests that many current developments may soon be outdated due to rapid advancements in model capabilities.
Reinforcement Fine Tuning (RFT) is highlighted as a promising area for specialization in models. Developers are encouraged to create environments that reward models for RFT, with a discussion on whether to rush into fine-tuning or wait for future advancements. Speculation arises regarding past attempts to integrate reinforcement learning (RL) and reasoning in language models, with the belief that models that think before acting show improved performance.
In a discussion with Ilya, the speaker conveys that achieving human-level general intelligence (HGI) is distant without a general reasoning paradigm in language models. The evolution of research at OpenAI is noted, with a resurgence in RL research leading to the development of a reasoning paradigm. Internal disagreements within OpenAI are acknowledged, particularly regarding the balance between pre-training and the need for new research into RL.
OpenAI has successfully scaled the pre-training paradigm while acknowledging the necessity for further research directions. Researchers debated the importance of reasoning and RL in enhancing data efficiency, believing that data limitations would arise before compute limits. The challenge lies in improving algorithmic data efficiency while scaling compute resources.
The discussion also touches on the coding capabilities of models like Codex, with personal experiences emphasizing their effectiveness and the learning curve regarding their limitations. Both speakers acknowledge the rapid progress in AI while recognizing constraints, such as GPU limitations and the need for advancements in the software development lifecycle.
Advancements in AI are expected to impact various remote work tasks, particularly in roles like virtual assistants. Understanding AI's capabilities and limitations is essential for anyone in remote work. The alignment of AI with user goals is critical, distinguishing between safety alignment to prevent harm and instruction-following alignment to execute commands.
Noam Brown discusses the scaling of test time compute to enhance AI problem-solving capabilities. His team is investigating both collaborative and competitive dynamics within multi-agent systems, arguing that human intelligence has historically thrived on cooperation and competition. This suggests that AIs could similarly evolve and generate advanced solutions through collective efforts.
Brown compares his multi-agent approach to Jim Fenn's Voyager skill library, asserting that their methodology diverges from traditional practices in the field. He critiques the multi-agent domain for its misguided heuristic approaches, emphasizing the importance of scaling in research. In discussing poker strategy, he highlights the distinction between Game Theory Optimal (GTO) strategies and exploitative play, acknowledging the inherent trade-off between these strategies.
With a background in AI for poker, Brown explains that while AIs can execute unbeatable strategies against both skilled and unskilled players, they struggle against weaker opponents due to their lack of adaptability. Transitioning from poker to diplomacy, he points out that game theory does not translate well to collaborative environments, where understanding and adapting to other players is vital.
Brown emphasizes the potential for techniques developed in diplomacy to enhance exploitative strategies in poker AI, advocating for further research in this area. He notes that current poker AIs primarily utilize pre-computed GTO strategies and lack adaptability to player behavior. The conversation also touches on the concept of world models in AI, suggesting that as models scale, they implicitly create a world model without needing explicit modeling.
The speaker discusses the evolution of world models in AI, suggesting that advanced models may develop a theory of mind, understanding other agents' actions and motives implicitly. They emphasize the "bitter lesson" in AI, which highlights the importance of self-play in achieving better results than human training, as demonstrated by systems like AlphaZero.
The speaker notes challenges in optimizing self-play objectives outside of two-player, zero-sum games, suggesting that the analogy with AlphaGo may not easily translate to other contexts. They express admiration for generative media, highlighting the significance of autoregressive emission and the excitement surrounding image and video generation.
In discussing robotics, the speaker reflects on their master's in robotics and the slow research cycle associated with physical hardware. They express a neutral stance on humanoid robots, acknowledging the value of non-humanoid designs like drones. The speaker questions the limitations of humanoid shapes and the potential eeriness of unconventional designs.
The conversation touches on the qualities that make a good hire, including skills in deception and detecting deception, and reflects on the nature of games with perfect versus imperfect information. Noam Brown shares his expertise in AI and frame-frame information games, particularly in poker, where hidden information is limited. He explains the complexity of scaling hidden possibilities, especially in games like Omaha Poker and Stratego.
Noam discusses the potential for developing superhuman bots for games like Magic the Gathering, emphasizing the need for general reasoning techniques in AI advancements for complex games.
This summary was generated from the episode transcript and can contain mistakes.