#244 - GPT-5.5 Instant, Grok 4.3, OpenAI vs Musk
Monday, 11 May 2026 · 4 min read · Listen to the episode ↗
In episode 244, the hosts dig into OpenAI's release of GPT-5.5 Instant as the new default ChatGPT model, which scores 81 percent on AME competition math versus 65 percent for its predecessor and carries the first high cyber risk rating under OpenAI's preparedness framework, though that rating reflects maximum reasoning effort rather than the low effort at which the model is actually deployed.
OpenAI released GPT-5.5 Instant as the new default ChatGPT model, scoring 81 percent on AME competition math versus 65 percent for GPT-5.3 Instant and 85 percent on GPQA versus 78 percent. Despite being lighter-weight than GPT-5.5 Thinking, it outperforms that model on Capture the Flag and CVE Bench cyber benchmarks and is the first instant model rated high cyber risk under OpenAI's preparedness framework. That rating reflects maximum elicited capability at high reasoning effort, not the low reasoning effort at which it is actually deployed. Self-improvement evaluations were skipped on the grounds that it is less capable than 5.5 Thinking, a decision one host argued sets a bad precedent. Synthetic biology evaluations showed regressions, producing a mixed rather than clean uplift picture.
OpenAI published a post-mortem on the goblin word frequency problem, tracing it to a reinforcement learning reward signal instructing playful language that the training loop interpreted as creature metaphors. A nerdy persona variant representing roughly 2 to 3 percent of ChatGPT responses produced two thirds of all goblin mentions and saw a 4,000 percent increase. The behavior first appeared in GPT-5.1 around November 2025 and compounded through subsequent versions because training runs for GPT-5.2 through 5.4 had dependencies on outputs from earlier decimal versions, meaning OpenAI cannot simply revert to a prior checkpoint without redoing a substantial chain of training. A data feedback loop where deployed model outputs feed back into supervised fine-tuning was identified as the self-reinforcing mechanism.
XAI launched Grok 4.3 on the API before publishing any announcement, with a 40 percent reduction in input costs, 60 percent reduction in output costs, a 1 million token context window, always-on reasoning, and roughly 100 tokens per second throughput. One host characterized it as comparable to Claude 4.6 or recent GPT instances, smart and cheap but not a frontier model. Initial hard-coded reasoning caused user complaints including analysis paralysis and freezing behavior, with configurable reasoning added afterward in what one host described as a rushed release. Grok 4.4 is expected within weeks and Grok 5 is targeted at 10 trillion parameters later in the year.
Mistral released Medium 3.5, a 128 billion parameter dense model with a 256,000 token context window scoring approximately 78 on SWE-bench verified. Its dense architecture contrasts with the mixture-of-experts direction taken by DeepSeek and Qwen, trading lower inference efficiency for simpler deployment. Mistral replaced its Apache 2.0 license with a modified MIT license restricting free use for high-revenue companies, interpreted as preventing well-resourced competitors from using the model without payment. Pricing at 1.50 dollars per million input tokens is cheap relative to Claude but expensive relative to open source models with a comparable benchmark profile.
Anthropic updated Claude managed agents with three features. Dreaming is a scheduled background process that reviews past sessions and memory to self-improve agents over time, using spare GPU capacity during low-demand periods. Outcomes lets users define success criteria evaluated by a separate grader agent in its own context window, creating an actor-critic loop independent of the worker agent's reasoning. Multi-Agent Orchestration allows a lead agent to delegate to specialist sub-agents with their own models, prompts, and tools.
Anthropic signed a deal with SpaceX to access the Colossus One data center, gaining over 300 megawatts and more than 220,000 Nvidia GPUs expected online within one month. The deal is notable because Musk has publicly criticized Anthropic and xAI competes directly with it. xAI was reportedly running at approximately 11 percent GPU utilization versus a typical 30 to 40 percent, implying the majority of tens of billions in capital cost was going unused, which one speaker argues makes competing as a data center and cloud business more strategically sensible for xAI than competing as a frontier model lab. Anthropic is in talks to raise funds at a 900 billion dollar valuation, exceeding OpenAI's most recent 852 billion dollar valuation and more than doubling Anthropic's February valuation of 380 billion dollars. Dario Amodei cited an 80x revenue jump in the first quarter of this year. Anthropic announced a 1.5 billion dollar joint venture with Blackstone, Hellman and Friedman, and Goldman Sachs targeting enterprise AI services, with Bloomberg reporting OpenAI is raising for a similar venture at a larger scale. Both structures direct asset managers to push portfolio companies toward adoption, functioning as a shortcut to mid-market enterprise penetration using a Palantir-style embedded engineer model. Anthropic's deal with FIS, which processes nearly 12 percent of the global economy targeting money laundering workflows, caused FIS shares to jump 7 percent.
The OpenAI versus Elon Musk trial began with jury selection April 28th. Musk dropped fraud charges just before trial started and observers consider his remaining legal case relatively weak. Text messages show board member Siobhan Zilis, who is also the mother of four of Musk's children, asking Musk whether to distance herself from Altman, with Musk instructing her to stay friendly and collect information. Zilis testified Musk wanted OpenAI to merge into Tesla and offered Altman a Tesla board seat, undercutting Musk's framing as a guardian of a charitable mission. Greg Brockman's personal diary contains a note asking what it would take to reach one billion dollars in net worth, contradicting his court testimony emphasizing mission. Musk admitted xAI used OpenAI models to generate training data, calling it industry standard, which the hosts note contradicts the framing applied to Chinese model providers doing the same with Claude.
This summary was generated from the episode transcript and can contain mistakes.