Sam Lehman: What the Reinforcement Learning Renaissance Means for Decentralized AI
Wednesday, 30 April 2025 · 3 min read · Listen to the episode ↗
Sam Lehman explores the evolution of AI scaling through three phases, emphasizing reinforcement learning's (RL) critical role in enhancing large language models (LLMs). He discusses the DeepSeq moment, which uses innovative RL techniques to improve reasoning capabilities, and highlights the importance of decentralized AI networks for fostering collaboration and creativity over centralized models. Lastly, he addresses concerns around the reliance on human data and the need for synthetic data to optimize AI performance, paving the way for future advancements in decentralized AI.
Sam Lehman discusses the evolution of AI scaling, outlining three key phases. The first phase focused on pre-training, where increasing data and compute improved model performance. Early researchers at OpenAI and DeepMind established scaling laws, revealing that models initially had too many parameters relative to the data, necessitating a balance of data and compute. The second phase introduced inference time compute scaling, emphasizing the importance of allowing models more time to think, which led to the development of reasoning models. These models demonstrated that smaller architectures could outperform larger ones by utilizing longer inference times, as highlighted in Google DeepMind's "Scalene LLM test time compute" paper.
The conversation transitions to the third phase, which centers on the role of reinforcement learning (RL) in enhancing large language models (LLMs). The DeepSeq moment marked a significant innovation in RL application, employing a performant base model and a novel RL process using the GRPO algorithm to enhance reasoning capabilities. Following DeepSeq's introduction, there was a surge in the exploration of new RL algorithms within the industry. Lehman details the DeepSeq stack, which includes a 671 billion parameter base model (DeepSeq v3) and two subsequent models, R10 and R1, created using different RL processes. R10 was trained primarily on math and coding questions, learning through trial and error with binary rewards, which allowed it to improve problem-solving skills without relying on curated human data.
The complexities of language processing are discussed, particularly the challenges of thinking in multiple languages. Concerns are raised about the impact of human data on RL, questioning whether it stifles model creativity and performance. The history of AlphaGo and AlphaZero is referenced, noting that removing human data allowed the models to learn independently, leading to remarkable outcomes. The importance of allowing AI models to explore beyond human preferences to foster creativity is emphasized, alongside the challenges of applying RL techniques to non-verifiable domains, such as creativity and writing.
The discussion includes the importance of high-quality question and answer pairs for initial training and the generation of synthetic data to enhance model performance. Future challenges involve reducing reliance on high-quality human data to create more efficient models. The concept of a decentralized RL network is introduced, comprising a foundation model, a gym for diverse reasoning behaviors, and a refinery for optimization. Diverse environments are advocated to facilitate RL across various domains, emphasizing the need for robust verifiers to assess model outputs.
The benefits of decentralization over centralization are discussed, with an analogy comparing closed-source labs to an isolated smart child versus an open school that fosters collaboration. The importance of decentralized AI is emphasized, expressing concerns about relying on centralized companies for problem-solving. The potential for monetizing contributions, such as generating reasoning traces or creating environments that improve model performance, is explored. The idea of a global decentralized network for crowdsourcing tasks and verified reasoning traces is also discussed.
The concept of a global open-source decentralized world model is introduced, supporting continuous improvement through ongoing data generation. The conversation shifts to modular sparse models, which allow for specialized training in narrow domains, creating a supermodel through plug-and-play experts. Future writing plans include a focus on modular mixture of experts models, emphasizing the architecture that activates specialized parameters during inference. The necessity of simultaneous training for innovation through specialized sub-experts is highlighted.
Concerns arise about whether decentralized models can compete with those from established Frontier Labs, while optimism about open-source AI is expressed, stressing the advantages of global access and ownership. The dominance of major players like OpenAI is noted, which creates barriers to entry and fosters user loyalty. The issue of user lock-in, particularly regarding memory features in AI, is reiterated as a significant challenge.
Lehman shares insights on recent discussions about RL, clarifying that a paper suggested minimal performance differences between base models and reasoning models. While reasoning models tend to provide correct answers on the first attempt, base models may require multiple tries. He disputes claims that "RL is dead," asserting its necessity for eliciting desired reasoning behavior in models. He highlights the importance of performant base models for effective RL, concluding that both pre-training and RL should work in tandem for optimal AI development.
This summary was generated from the episode transcript and can contain mistakes.