#235 - Sonnet 4.6, Deep-thinking tokens, Anthropic vs Pentagon
Tuesday, 3 March 2026 · 4 min read · Listen to the episode ↗
The podcast discusses notable advancements in AI, focusing on Sonnet 4.6's enhanced capabilities in reinforcement learning and high-dimensional challenges for large language models (LLMs). The rising significance of "deep-thinking tokens" in evaluating AI reasoning is explored, alongside policy concerns regarding Anthropic's compliance with government regulations, particularly in relation to Pentagon demands and security impacts. Additionally, developments in cryptocurrencies and blockchain are briefly alluded to in the context of AI's growing influence on tech markets.
Andrei Kerenkov and Jeremy Harris discuss notable AI developments, focusing on updates like Sonnet 4.6, which features significant improvements such as an increase in contact size to 1 million, reflecting advancements in post-training reinforcement learning. The ARKGI 2 benchmark achieved a score of 60.4%, indicating strong performance but still trailing behind Opus 4.6 and Gemini 3D. This benchmark evaluates AI's human-level performance and highlights challenges for large language models (LLMs) in high-dimensional visual problems.
Francois Chedet, the researcher behind ArcGi benchmarks, anticipates reaching ArcGi 4 before achieving Artificial Superintelligence (ASI). The design of the benchmark emphasizes generalization from limited data, requiring model submissions for scoring. Ryo notes a lack of a clear leader in LLMs, despite companies using their own models, and mentions a performance uplift from 5% a year ago to 20% now, indicating a potential phase transition.
Recent updates include Google’s Gemini 3.1 Pro, which achieved 77.1 on ArcGi 2, featuring interactive capabilities like generating a 3D starling murmuration. The podcast discusses the pricing of AI models, highlighting that Claude is relatively affordable compared to Claude Opus 4.6, which is significantly more expensive. Anthropic maintains premium pricing, appealing to enterprise customers, but concerns about commoditization arise as models become more similar.
The significance of ROCK 4.20, currently in public beta, is debated, with Elon Musk claiming it will be smarter and faster than its predecessor. Changes at XAI include its integration into SpaceX, with some co-founders leaving. The model features four agents that engage in debate and fact-checking, raising liability concerns when users upload medical data for second opinions.
Perplexity has introduced an AI agent capable of assigning tasks to other agents, while Meta's partnership with AMD involves a significant investment in chips, focusing on AMD's MI540 GPUs and CPUs. Metax, a challenger to Nvidia, has raised $500 million to develop processors aimed at outperforming Nvidia GPUs for LLM training and inference.
World Labs has raised $1 billion to develop world models, with their first product, Marble, enabling the creation of editable 3D environments. The startup Simile has secured $100 million to predict human behavior and simulate actions for training AI agents. OpenAI faces challenges with delays in its Stargate AI data centers due to disagreements with Oracle and Softbank over control.
China aims to increase its leading-edge chip output significantly, focusing on scaling existing capabilities rather than advancing to cutting-edge technology. Research advancements in adaptive optimizers for neural networks are discussed, highlighting new methods that show promise in improving performance by selectively masking updates.
A new paper titled "Think Deep, Not Just Long" examines how to evaluate LLM reasoning through "deep thinking tokens," which fluctuate until later layers of the neural network. The concept of "deep thinking layers" is explored, noting that more complex tokens require deeper processing. The conversation also touches on token count, where an initial increase in tokens is linked to better performance, but excessive tokens can lead to reduced accuracy.
The discussion highlights the significance of consistent signals in decision-making and introduces the idea of a cautious optimizer, which favors stable updates to minimize noise. A new safety benchmark, Nessie, aims to identify errors in AI models, establishing minimum performance standards. Analysis from Epoch AI suggests that rapid AI progress is primarily driven by improved data and optimization techniques rather than theoretical breakthroughs.
The conversation transitions to policy and safety issues, with Anthropic's CEO asserting that Pentagon threats do not influence their AI usage stance. Anthropic maintains strict limitations against using their technology for spying on U.S. citizens or lethal autonomous weapons. The potential application of the Defense Production Act to compel Anthropic to produce goods for national defense raises questions about the government's contradictory stance.
The Pentagon has secured a deal to utilize GROC in classified systems, providing an alternative to Anthropic. Employee pressure is influencing decisions at Anthropic, where staff are celebrating recent developments that coincide with a relaxation of the company's safety and security commitments. Anthropic has released a report on distillation attacks, revealing significant concerns about export control policies and the risks of AI security.
OpenAI's initiatives to combat malicious AI uses are highlighted, including a monthly report on misuse cases. The conversation concludes with an encouragement for listeners to engage with the podcast and stay informed about AI developments.
This summary was generated from the episode transcript and can contain mistakes.