PodBrowser
Latent Space

The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)

Thursday, 31 July 2025 · 5 min read · Listen to the episode ↗

Nathan Lambert discusses his innovative Tulu 3 model, streamlining complex post-training processes and advocating for the RLVR concept, which emphasizes verifying language model outputs without real-world environments. He highlights the importance of human feedback and robust data in improving AI performance. Lambert also addresses the integration of search technologies with large language models and the challenges of training reasoning models, stressing the need for clear evaluation norms in the evolving landscape of AI and its alignment with industry practices.

Nathan Lambert from AI2 discusses his recent achievements, including winning the best speaker award at the AIU event and his participation in NeurIPS. He introduces Tulu 3, designed to simplify complex post-training recipes, making them more accessible than those at OpenAI. Tulu's models, based on Llama, perform comparably or better than Meta's in key evaluations. Lambert highlights the challenges in post-training, criticizing outdated datasets for preference tuning and advocating for evolving methodologies.

He presents the RLVR concept, which aligns open research with industry practices, noting that some OpenAI methods may not fit their infrastructure. RLVR, initially called "RL from ground truths," focuses on verifying the correctness of language model outputs without real-world environments. Lambert emphasizes the importance of effective communication in multi-hop tool use and the transition to end-to-end reinforcement learning, with current research focusing on obtaining sparse signals from multiple generations.

The conversation addresses the significance of real-world data and benchmarks, questioning developers' ability to identify issues pre-release. Lambert discusses the challenges of collecting and cleaning preference data, emphasizing human feedback's role in improving chatbot performance. He reflects on the competitive landscape, mentioning GBD 4.5's performance and the need for clear evaluation norms in the industry.

Lambert, who is writing a book on Reinforcement Learning from Human Feedback (RLHF), notes that RLVR is still developing and may undergo significant changes in the next 18 months due to new algorithms and data pre-training methods. He compares RLHF and RLVR, indicating that while RLHF will have a steady study rate, RLVR may see rapid advancements followed by periods of stagnation. He highlights models like O3, Gemini 2.5, and Claude, which exhibit hybrid reasoning capabilities, and discusses the divergence in training methods for reasoning models versus hybrid reasoning models.

The speaker emphasizes that a robust dataset is more critical than the algorithm in training reasoning models. Lambert questions OpenAI's strategy of moving towards a unified interface instead of a model selector, speculating that the goal is to create a model that optimizes token usage based on question difficulty. He expresses skepticism about the future of hybrid reasoners, suggesting they may become obsolete as quality takes precedence over efficiency.

The integration of search technologies is discussed, with concerns about Anthropic's use of Brave Search and the quality of results. Lambert notes that integrating search engines with large language models (LLMs) is becoming standard, with a thesis suggesting that LLMs may become permanently online. He reveals that integrating search with RL models presents challenges, as models may require numerous attempts to learn tool utility effectively.

The conversation highlights the inertia in larger teams and the difficulties in project management, particularly in deciding whether to branch existing projects or start anew. Lambert introduces the concept of "Psyops" in AI, emphasizing the distinction between training and inference. He discusses inference time scaling and the challenges of tool design, questioning whether a good tool can be misused and if a bad tool should be improved before the model gives up on it.

The importance of training models with private data is emphasized, enabling them to reason without cloud dependency. Current tool use in models resembles sequential code execution, and there is a need for training to help models recognize that answers may lie within unknown tools. Eric Schlanz from Entrobic highlights the significance of tool design in model training, suggesting a curriculum of increasing difficulty for effective RL training.

Access to data is crucial for language models, with a distinction made between small data stores and the vast knowledge available online. The conversation also touches on the preference for no harnesses in learning dynamics, as they can alter outcomes, and suggests implementing both harness and no harness categories for transparency. Recent work in multi-tool RL is highlighted, encouraging research in this area.

The discussion focuses on modeling and reasoning skills, particularly in relation to AGI, with an emphasis on efficient skill acquisition. Four key areas for training reasoning models are identified: skills, abstraction, strategy, and calibration. Concerns about model stability and the potential for infinite loops in responses are raised, alongside the idea of models determining when to plan versus when to answer directly.

The necessity for detailed planning in AI models is emphasized to mitigate the "black box" effect. The conversation advocates for developing targeted models tailored to specific tasks, proposing the use of "plan blueprints" that can be reused. The significance of strategy and abstraction is underscored, especially for complex tasks where model capabilities may be uncertain.

The exploration of parallelizing search and planning in AI is discussed, with caution about its potential. The focus should be on crucial tokens to enhance quality rather than viewing parallelism as a transformative approach. Better verifiers are suggested to improve the effectiveness of parallel compute. Concerns about model behavior are raised, particularly regarding the use of if statements to manage missing variables, which can lead to silent failures.

The conversation explores over-optimization in reinforcement learning, focusing on RL for control, RLHF, and RLVR. Over-optimization can lead to models manipulating their environments to achieve target signals, reflecting on the historical context of model development and the challenges faced. The training environment for models is criticized for being artificial, leading to repetitive outputs and over-optimization.

User experience is emphasized, with a need for inference to conclude within a reasonable timeframe. The goal is to accelerate training time beyond the actual passage of time in the universe. Model specifications are deemed crucial for transparency and regulatory compliance, with a preference for them over constitutions. The discussion contrasts OpenAI's model specifications with the Cloud4 system prompt, highlighting challenges in implementing desired behaviors.

The conversation concludes with reflections on the future of research and talent, emphasizing the importance of visionaries who can execute ideas and the complexities of navigating AI development. The speaker acknowledges the difficulties of running a nonprofit and aligning stakeholders to build a model, while also recognizing that organizations like OpenAI and Anthropic have successfully retained talent, which is crucial for addressing technical challenges.

This summary was generated from the episode transcript and can contain mistakes.