#242 - ChatGPT Images 2.0, Qwen 3.6 Max, Kimi-K2.6
Wednesday, 29 April 2026 · 4 min read · Listen to the episode ↗
ChatGPT Images 2.0 leads the episode, described as the largest image generation leap since DALL-E 3, using a transformer token-based architecture rather than a diffusion model and demonstrating precise text rendering, working SVG code, and accurate GUI screenshots. Kimi K2.6 from Moonshot AI matches GPT-4.5 benchmarks with one trillion total parameters but only 32 billion active per inference cycle, trained natively in int4 quantization to optimize deployability under export control constraints.
ChatGPT Images 2.0 is described as the biggest image generation leap since DALL-E 3, based on ELO human preference rankings. It uses a transformer token-based architecture rather than a diffusion model, aligning the image pipeline with OpenAI's existing LLM infrastructure to avoid maintaining two separate systems. The model can generate precise text, working SVG code, and accurate screenshots of desktop GUI applications, with training data appearing to include screenshots, Adobe software interfaces, and Starcraft imagery. It also shows improved stylistic range across photo realism, gritty 90s aesthetics, fashion book styles, and comic book styles, addressing the smooth sheen that previously made OpenAI images identifiable. OpenAI declined to confirm architectural details at a press briefing, leaving chain-of-thought and reasoning claims unconfirmed. Jeremy characterizes OpenAI as holding a robust advantage over Anthropic and arguably Google in multimodality.
Kimi K2.6 from Moonshot AI benchmarks on par with GPT-4.5 and comparably to Anthropic's top models. It has 1 trillion total parameters across 384 experts in a mixture-of-experts architecture, with only 32 billion active parameters per inference cycle. It uses multi-head latent attention to compress query and key vectors and reduce memory overhead, a technique described as part of DeepSeek's legacy and especially important for Chinese labs facing export controls. Kimi K2.6 is trained natively in int4 quantization rather than post-training quantized, sacrificing some precision to optimize deployability from the start, and supports long-horizon multi-day agentic tasks with parallel agent spawning.
Qwen 3.6 Max Preview is not open source and is available only via API, a departure from previous Qwen releases. Qwen has overtaken Meta's Llama as the most deployed self-hosted model in the world. Alibaba shut down the free tier of Qwen Code just days before the release, suggesting the free tier was used to build network effects that Alibaba is now monetizing. Qwen 3.6 Max Preview benchmarks approximately on par with Claude 4.5 Opus, though a caveat is raised that proprietary models tend to outperform benchmark scores relative to open source models, making it difficult to determine how many months behind Chinese labs actually are.
Google's Deep Research Max, built on Gemini 3.1 Pro, scores approximately 86 percent on the Browse Comp benchmark while GPT-5.4 scores approximately 59 percent, with the non-Max Deep Research version comparable to GPT-5.4. On Humanities Last Exam, Gemini Deep Research and GPT-5.4 are essentially tied within error bars. Deep Research now supports the Model Context Protocol, allowing access to proprietary data sources not accessible on the web. In deep research and report generation, OpenAI GPT-4.5 is considered Google's primary competition, not Anthropic.
SpaceX is working with Cursor and holds an option to buy the startup for 60 billion dollars, with SpaceX paying either 10 billion dollars for collaboration or exercising the acquisition option later in the year. The deal is framed as helping train better coding models for XAI. Context flagged includes Grok 4.2 appearing at the bottom of recent coding benchmarks, a promised Elon Musk coding model that never materialized, nearly the entire founding team of approximately 12 people having left XAI, and Cursor losing market share to Claude Code. A caveat is raised that Cursor only fine-tunes rather than training foundation models from scratch, which is a different capability from what XAI may need.
Cerebras Systems is planning a mid-May IPO at a 23 billion dollar valuation, reporting 510 million dollars in top-line revenue for 2025 but a net loss of 75 million dollars when one-time items are excluded, making the non-GAAP numbers worse than GAAP, which is flagged as unusual. Customer concentration risk is flagged because its two largest customers are OpenAI and AWS. OpenAI building its own inference chips and AWS pushing Trainium are cited as potential headwinds.
Anthropic is receiving 5 billion dollars from Amazon and has pledged 100 billion dollars in cloud spending over 10 years, a roughly 20-to-one ratio of commitment to investment. Mozilla used Anthropic's Mythos to find and fix 271 bugs in Firefox, with Firefox's CTO describing it as a transitory moment requiring a one-time overhaul of all software to surface latent vulnerabilities. The NSA is reportedly using Claude MFOS despite a DoD blacklist designating Anthropic as a supply chain risk, and a separate unauthorized Discord group gained access through a third-party provider, representing a second Anthropic security breach. The conflict between DoD policy and NSA use is predicted to be difficult to sustain and may ultimately go to court.
Deezer reports that 44 percent of songs uploaded to its platform daily are AI generated, growing from 10,000 per day in January 2024 to 60,000 per day by January 2025, though AI tracks account for only one to three percent of total streams and 85 percent of those streams are flagged as fraudulent and demonetized. Meta is planning to lay off approximately 8,000 employees globally in 2025, representing roughly 10 percent of its workforce, and has a mandatory program recording employee mouse movements, keystrokes, and screen content on work laptops with no opt-out option, interpreted as training AI to perform employees' jobs.
This summary was generated from the episode transcript and can contain mistakes.