PodBrowser
Last Week in AI

#242 - ChatGPT Images 2.0, Qwen 3.6 Max, Kimi-K2.6

Wednesday, 29 April 2026 · 4 min read · Listen to the episode ↗

In episode #242, Andre Karankov and Jeremy Harris highlight advancements in AI, focusing on ChatGPT's image generation model and Alibaba's Qwen 3.6 Max. They discuss improvements in realism and diverse styles in new models, though express confusion over the necessity of new releases. Kimi K 2.6 by Moonshot AI is noted for its substantial parameter count and sparse architecture, reflecting ongoing innovation in AI capabilities. Additionally, developments in AI governance and implications for national security are explored.

Andre Karankov and Jeremy Harris discuss significant advancements in AI, particularly focusing on ChatGPT's new image generation model and developments in Chinese AI models. They highlight ChatGPT's capabilities in generating precise text and complex SVG code, showcasing improvements in desktop GUI image generation. Jeremy draws parallels between nuclear treaties and AI development, suggesting that limitations in AI may be more symbolic than substantive without a fundamental shift in compliance incentives.

The hosts express confusion about the need for new image generation models, acknowledging improvements in reasoning and code generation but lamenting the lack of technical details. They note that newer models exhibit diverse styles and improved realism, with applications in creating infographics and posters. The conversation includes Alibaba's Qwen 3.6 Max Preview, which, while not open source, represents a significant improvement over earlier models but still falls short of competing with OpenAI.

They discuss the rapid pace of LLM releases from both Western and Chinese developers, with proprietary models often outperforming open-source ones. The hosts compare current AI model benchmarks, noting some models lag behind the frontier by five to six months. Google has introduced "Deep Research" and "Deep Research Max" agents based on Gemini 3.1 Pro, enhancing research capabilities with improved extended test time compute. Benchmark scores reveal that Google's Deep Research excels in deep search QA and niche fact-finding tasks.

The asynchronous workflow of Deep Research Max enhances user experience, and advancements in test time compute lead to qualitatively different interactions with these models. Mozilla's use of the tool Mythos for bug detection in Firefox is noted, emphasizing its role in improving software security. Concerns about software security are raised, particularly regarding the implications for national security.

The Starbucks ChatGPT app is mentioned, allowing users to place orders via ChatGPT, though the experience is described as clunky. SpaceX is considering acquiring Cursor for $60 billion or collaborating for $10 billion, with Cursor expected to assist in training better coding models for XAI. However, concerns arise about Cursor's ability to train frontier models, as they primarily fine-tune existing models.

XAI faces challenges, having not yet released a true frontier model, and the collaboration with Cursor may not address its foundational training needs. The recent talent drain at Cursor complicates the situation further. The conversation also touches on Cerebus, an AI chip startup that has filed for an IPO, known for its unique chips designed for AI inference.

Recent fundraising highlights include a startup raising $180 million for biological-based AI models and Core Automation seeking $500 million to $1 billion for data-efficient models. Anthropic has secured $5 billion from Amazon, reflecting a long-term cloud lock-in deal. The conversation emphasizes the significance of including individuals who can enhance capital in founding teams, particularly in the context of ongoing workforce reductions at Meta.

In the semiconductor sector, China's dual strategy of advancing domestic fabrication while increasing imports of US chip-making equipment is discussed. Google is exploring new chip designs to enhance AI performance, including a memory processing unit and an inference-optimized TPU. The release of new models, such as Kimi K 2.6 by Moonshot AI, showcases ongoing innovation in AI, with Kimi K 2.6 featuring one trillion parameters and a sparse architecture.

A paper titled "Infusion Shaping Model Behavior by Editing Training Data via Influence Functions" explores how to influence model behavior by identifying and perturbing training documents. Recent updates on MIFOS indicate that the NSA is reportedly using it despite being on the DoD blacklist due to supply chain risks. The integration of frontier labs into national security is seen as problematic, with cybersecurity for AI models presenting new challenges.

Research collaborations, such as one between the University of California, San Diego, and Together AI, focus on stable looped language models, resulting in a reduction in validation perplexity. The conversation emphasizes memory management in transformer models, discussing the need to reduce memory usage while enhancing computational efficiency.

The introduction of OcuBench evaluates AI agents in real-world professional tasks through simulated environments, covering 100 scenarios across 65 domains. In the realm of synthetic media, Deezer reports a rise in AI-generated music, although it accounts for only a small percentage of total streams. Concerns about deepfakes are raised, with celebrities able to request the removal of AI-generated deepfakes on platforms like YouTube.

This summary was generated from the episode transcript and can contain mistakes.