PodBrowser
Last Week in AI

#235 - Opus 4.6, GPT-5.3-codex, Seedance 2.0, GLM-5

Monday, 16 February 2026 · 4 min read · Listen to the episode ↗

In episode #235, Kurenkov and Harris analyze notable AI advancements, highlighting the releases of Opus 4.6 and GPT-5.3 Codex, both of which demonstrate significant improvements in speed and capabilities. They discuss Seedance 2.0, a text-to-video model expanding creative potential, and GLM-5, which features enhanced performance metrics and adaptability. The conversation emphasizes the competitive landscape among AI developers, revealing the critical role of innovative models in reshaping productivity and ushering in advancements in tasks across sectors, including video generation and coding.

Andrei Kurenkov and Jeremy Harris discuss notable AI developments, focusing on recent model releases such as Opus 4.6, GPT-5.3 Codex, and advancements from Chinese companies. They emphasize a shift towards research and progress in AI, moving away from business news. Anthropic's Opus 4.6 introduces agent teams for task distribution, featuring a 1 million token context window and a version that is 2.5 times quicker than its predecessor. Initial feedback highlights significant improvements in capabilities, particularly in vibe coding and task decomposition into parallel workflows. This release reflects competitive pressure in the AI space, evolving from a developer tool to a universal knowledge worker.

OpenAI's GPT-5.3 Codex is noted for its 25% speed improvement and advanced capabilities, including its role as the first high-capability model for cybersecurity. The hosts discuss the challenges of evaluating AI models, as their behavior may change based on evaluation awareness, and the complexities in measuring AI self-improvement. Significant advancements in AI capabilities are acknowledged, particularly with Codex 5.3 achieving a 77.3% performance on the Terminal benchmark. OpenAI's collaboration with Cerebras aims to diversify hardware options, with Cerebras handling low-latency use cases while Nvidia supports high-capability models.

The podcast critiques the concept of recursive self-improvement, suggesting that AI primarily enhances productivity rather than creating smarter models. A data scientist utilized GPT-5.3 codecs to create new data pipelines and enhance visualization, emphasizing the complexity of understanding model performance through qualitative experiences rather than benchmark scores. The integration of AI in software companies is deemed crucial, with those not adopting AI seen as lagging behind.

Codex Spark's performance is highlighted, achieving over 1000 tokens per second, although trade-offs exist, including rate limits and a 128 context window. Speculation arises regarding the architectural implications of deploying Codex Spark on various hardware. The competition between Cerebras and Nvidia is noted, particularly Nvidia's high profit margins on GPUs. Codex Spark is distinct from Codex 5.3, optimized for fast inference with a 59% score on the Terminal benchmark.

Google's Gemini Free DeepThink update achieved an 84.6% pass rate on the ArcAGI2 benchmark, surpassing Opus 4.6's 68.8%. However, the lack of a system card for the new model raises concerns about potential risks and the need for safety assessments in AI development. The introduction of Seedance 2.0, a text-to-video model, allows users to generate high-quality videos from text prompts, enhancing creative possibilities for agentic training. Anticipation grows for increased AI-generated video content as access to these models expands.

The conversation highlights significant advancements in AI models, particularly focusing on GLM 5 from JUPU AI, which boasts a model size twice that of GLM 4.5. GLM 5 features 744 billion parameters and a new RL framework called SLIME that enhances training efficiency. Deep Seek's new release offers a context window of 1 million tokens, greatly improving its utility for coding tasks. Concerns are raised about potential benchmark gaming and evaluation biases in comparing models like GLM 5 and Opus 4.6.

The discussion also touches on the versatility of GLM 5, which can be applied across various coding frameworks and non-coding tasks. The rapid advancements in AI, particularly in reinforcement learning, are noted, alongside the introduction of Cursor's Composer 1.5, which aims to compete with existing coding IDEs. The competitive landscape among LLM providers, including OpenAI, Anthropic, and Google, is discussed, with an emphasis on the increasing availability of open-source options.

The podcast discusses a recent funding round aimed at expanding into world model space, particularly in video generation and robotics. Aptronic, a company developing humanoid robots, raised $935 million at a $5.3 billion valuation, suggesting strong investor interest in humanoid robotics. A lighter topic arises with Anthropic's humorous Super Bowl advertisement targeting OpenAI, which elicited a defensive response from OpenAI leaders.

Waymo announced its sixth-generation car model is ready for passengers and high-volume production. In the realm of open-source projects, GLM5 is set to release on Hugging Face, while Quen3 Coder Next will provide an open-weight language model for coding agents. The training process for these models emphasizes distillation and supervised fine-tuning to enhance capabilities, though concerns about catastrophic forgetting persist.

Research advancements are discussed, particularly a paper on reasoning with 13 parameters, introducing LoRa (Low Rank Adaptation) as a method for efficiently adapting large models. The conversation also addresses the impact of task complexity on model performance, raising questions about the assumption that scaling AI will inherently enhance coherence. The concept of incoherence is introduced as a significant factor in AI behavior, suggesting a shift in focus towards reward hacking and goal mis-specification. The episode is noted for its dense and rapid pace, with hosts expressing appreciation for listener engagement.

This summary was generated from the episode transcript and can contain mistakes.