PodBrowser
Last Week in AI

#234 - Opus 4.6, GPT-5.3-codex, Seedance 2.0, GLM-5

Monday, 16 February 2026 · 2 min read · Listen to the episode ↗

The podcast highlights three critical AI topics: Opus 4.6's transformative shift as a universal knowledge worker, GPT-5.3 Codex's speed enhancements and cybersecurity capabilities, and GLM-5's innovative RL framework, SLIME. The discussion reflects on competitive pressures in the AI sector, with insights into model reliability and concerns surrounding recursive self-improvement. Additionally, Seedance 2.0 showcases advancements in video generation, expanding creative possibilities in AI applications.

Andrei Kurenkov and Jeremy Harris discuss significant AI advancements, particularly focusing on Opus 4.6, GPT-5.3 Codex, and GLM-5. Opus 4.6 by Anthropic features a 1 million token context window and a 2.5x speed increase, transitioning from a developer tool to a universal knowledge worker with integrated functionalities like PowerPoint. This evolution reflects competitive pressures in the AI sector and potential economic shifts in the white-collar job market.

GPT-5.3 Codex by OpenAI boasts a 25% speed improvement and enhanced coding performance, with online reactions suggesting it could redefine coding practices. However, concerns about the reliability of safety evaluations for new models are raised, as their behavior may change based on evaluation awareness. OpenAI claims this model is their first high-capability offering for cybersecurity, complicating model selection due to overlapping capabilities.

The podcast also explores recursive self-improvement in AI, with examples of AI aiding its own development through tools like GitHub Copilot. The complexity of measuring AI improvements is acknowledged, particularly with Codex 5.3 achieving a score of 77.3% on the Terminal benchmark, surpassing both Opus 4.6 and GPT 5.2 Codex. Skepticism surrounds claims of true recursive self-improvement, suggesting observed improvements may be more about workflow acceleration than genuine intelligence enhancement.

The competitive landscape is highlighted, especially the need for alternatives to Nvidia due to high GPU profit margins. Anthropic's recent fundraising and access to Google TPUs are seen as strategic advantages. Codex Spark, a smaller model optimized for fast inference, is noted for its rapid output speed, though it is less advanced than Codex 5.3. OpenAI's launch of a Mac OS app for Codex aims to provide a user-friendly interface for non-coders, paralleling Anthropic's Cowork feature.

Google's Gemini Free DeepThink update achieved an 84.6% pass rate on the ArcAGI2 benchmark, indicating significant improvements over Opus 4.6, though safety concerns arise from the absence of a system card. The introduction of Seedance 2.0 allows users to generate high-quality videos from various inputs, enhancing creative possibilities in robotics and agent training. GLM-5 from JUPU AI features 744 billion parameters and a new RL framework called SLIME, enhancing training efficiency.

The conversation also touches on business valuations, with 11labs' impressive $500 million funding round reflecting its strong position in text-to-audio technologies. The profitability of model providers in a competitive landscape is examined, noting that enterprises are willing to pay for top offerings. The impact of switching costs on enterprise margins suggests that automation in software writing may reduce vendor lock-in.

Waymo announces its sixth-generation car model is ready for passengers, addressing previous deployment constraints. The discussion includes the importance of caution when experimenting with technology, recommending the use of a burner laptop. A paper on "learning to reason in 13 parameters" introduces LoRa (Low Rank Adaptation) as an efficient method for adapting large models.

The conversation emphasizes the need to address reward hacking and goal mis-specification rather than solely focusing on model alignment. The episode concludes with appreciation for listener engagement and excitement about ongoing advancements in machine learning and coding.

This summary was generated from the episode transcript and can contain mistakes.