#230 - 2025 Retrospective, Nvidia buys Groq, GLM 4.7, METR
Wednesday, 7 January 2026 · 4 min read · Listen to the episode ↗
In episode #230, discussions center on the pivotal advancements in AI, including reasoning models and the acquisition of Groq by Nvidia, reflecting trends in talent acquisition and market strategy. The launch of GLM 4.7 showcases ongoing competition in AI coding models. Additionally, the potential for an AI bubble and the implications of data center investments are addressed, alongside rising concerns over alignment and safety protocols as major organizations, including OpenAI, implement new initiatives for AI governance.
Jeremy and Andrei reflect on key themes from 2025, highlighting advancements in reasoning models, agentic AI, and the rise of vibe coding. They discuss the increasing hype around world models and the rapid expansion of AI applications, particularly in the taxi sector. Significant releases in the first half of 2025 included R1 and Gemini 2.5, alongside Codex and Gemini CLI, while the second half saw a continued focus on reinforcement learning and improvements in terminal tools.
The conversation shifts to alignment and interpretability, with Google DeepMind and Anthropic exploring diverse strategies. Opening Eyes is attempting innovative approaches to reasoning without decoding tokens to reduce sycophancy. Concerns about the return on investment for large-scale AI experiments are raised, particularly regarding the sustainability of funding from sovereign wealth funds. The potential path to superintelligence by 2027-2028 is also mentioned, alongside a global focus on AI for national security and commercial interests.
Nvidia's acquisition of Grok is highlighted as a significant industry event, reflecting a trend where companies acquire talent through non-exclusive licensing to avoid antitrust issues. The year 2025 is characterized as pivotal for data centers, with rising concerns about an AI bubble as major labs invest heavily. Data center security is increasingly critical, especially regarding supply chain reliance and intellectual property protection.
Research advancements are noted, particularly in understanding toxic traits and misbehaviors in large language models (LLMs). The focus is shifting towards feature and control vectors rather than purely mechanistic interpretability. Activation steering techniques are gaining traction, representing a more fundamental approach to model interaction. A notable research paper from Anthropic on tracing LLM thoughts using feature vectors has provided improved insights into model operations, although it faced challenges in generalizability and computational expense.
Looking ahead to 2026, predictions include the emergence of on-device models for phones and laptops, driven by innovations like Google’s Gemini. There is an expectation for continued development of world models and realistic video generation. The conversation also explores the possibility of moving beyond the transformer architecture to hybrid or diffusion models, as current architectures may not be optimal. A key focus is on continual learning and its implications for AI research.
The discussion touches on the future of AI and data centers, emphasizing the necessity for data by the end of 2026 to justify ongoing investments. The term "Jagged Intelligence" is introduced, highlighting that while AI models excel in scientific tasks, they struggle with basic logic. Current advancements may improve performance, but there remains uncertainty about whether AI will achieve consistent intelligence and continuous learning akin to human common sense by 2026.
DeepSeek R1 is recognized as a pivotal paper of the year for its influence on AI research. OpenAI is consolidating its engineering, product, and research teams to enhance audio models for upcoming audio-first personal devices. Nvidia's acquisition of Grok for approximately $20 billion is characterized as unconventional, focusing on talent acquisition and technology licensing rather than a traditional buyout. This acquisition is pivotal for Nvidia's technical roadmap and market dominance in inference.
The conversation also covers the challenges faced by Chinese fabs in chip manufacturing, particularly their inability to purchase ASML's advanced EUV machines. Reports suggest that Chinese technicians are damaging machines while attempting to reverse engineer them. The discussion highlights U.S. efforts to restrict China's access to Nvidia technology, while Chinese companies like DeepSeek are reportedly improving their capabilities to produce chips that lag behind Nvidia's technology.
In open-source developments, z.ai has launched GLM 4.7, a new coding model that competes well with models like DeepSeq, although initial feedback has been mixed. Research advancements include the introduction of Democritus, a new paradigm for building causal models using large language models. The introduction of centered kernel nearest neighbor alignment (CKNNA) allows for comparison of how different models represent data, revealing that as models scale, their agreement on data representation increases.
The conversation also touches on AI safety legislation, with New York Governor Kathy Hochul signing the Raise Act, which requires large AI developers to disclose safety protocols and report incidents promptly. The speaker discusses their work on interoperability and activation oracles, introducing a method to use a model's internal outputs as inputs for another model. This approach aims to uncover hidden biases or intentions, allowing for monitoring of models for harmful goals.
Monitoring is crucial, with targeted follow-up questions enhancing the detection of model misbehavior. The conversation raises safety concerns, particularly in practical settings like health queries, where models may provide harmful responses. OpenAI's new initiatives include hiring a head of preparedness to oversee evaluations and safety mitigations related to advanced AI. The effectiveness of this new role will depend on control over safety measures.
This summary was generated from the episode transcript and can contain mistakes.