Grok 4.6 Shows How Fast Your AI Options Are Expanding
Thursday, 13 August 2026 · 4 min read · Listen to the episode ↗
Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, placing it ahead of Gemini K3 and tied with GPT-5.6 Sol at a price 60 percent cheaper than GPT-5.6 Sol and 73 percent cheaper than Fable 5, with one completed benchmark run recorded at 84 cents per task. Developer reactions split between those calling it a new default for time and cost efficiency and others warning it makes dangerous mistakes in security contexts.
Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, up five points from Grok 4.5, placing it ahead of Gemini K3, tied with GPT-5.6 Sol, and one to two points behind Fable 5 and Opus 5. On the GDPVAL benchmark measuring agent performance on economically valuable tasks, xAI claims Grok 4.6 overtook both GPT-5.6 Sol and Fable 5, though release benchmarks carry significant skepticism. The model is priced at two dollars per million input tokens and six dollars per million output tokens, making it 60 percent cheaper than GPT-5.6 Sol and 73 percent cheaper than Fable on a per-task basis, with Artificial Analysis recording a completed benchmark run at 84 cents per task. One developer testing it against 105 hidden bugs in two real repositories said it may become his default model for best combination of time, value, and cost, while another warned it makes dangerous mistakes in security work and gets defensive when challenged.
Elon Musk said Grok 4.7 has completed initial training, is significantly better than 4.6, and should be ready in three to four weeks, with supplemental training on a massive amount of SpaceX company data. Anthropic is reported to already have Fable 5.5 ready and waiting to be released, and Nathan Lambert noted the competitive vibe shifted from Anthropic being far ahead to model competition at all-time highs in roughly four weeks.
DeepSeek V4 Pro benchmarks leaked hours after Grok 4.6 launched, scoring 87.9 percent on Terminal Bench 2.1 and claiming to beat Fable by 0.2 percent on the CyberGym cybersecurity benchmark. However the Artificial Analysis benchmark gave it a score of only 53, just one point ahead of V4 Flash and behind Kimmy K3 and Mu Spark 1.2, raising concerns about benchmark maxing. One developer called it benchmark max slop that could not make a simple Minecraft clone. DeepSeek V4 Pro is priced at approximately one-twelfth the price of Fable, at 1.32 dollars per million input tokens and 3.96 dollars per million output tokens.
Fable 5 made up only 6 percent of tokens and 11.4 percent of dollars spent on Anthropic models, while GPT-5.6 Sol comprises 25 percent of OpenAI tokens and 23 percent of spend. Anthropic's 30-day data retention policy for US government safety checks is limiting enterprise adoption, as many businesses are unwilling to accept it. Ramp data suggests Fable 5 represents a new upper bound to what businesses are willing to pay, though that data carries selection bias from a cost-control-focused user base.
Cognition is in early talks for new funding at a 40 billion dollar valuation, up roughly 50 percent from the 26 billion dollar valuation it carried when it raised one billion dollars three months ago, and has doubled its revenue run rate to one billion dollars since that last round. One analyst predicted a hyperscaler could acquire Cognition for 60 to 100 billion dollars in stock within six to twelve months, and another argued Google should buy it for 200 billion dollars. Cognition's leadership said the company is not selling. Lovable announced a 400 million dollar Series C at a 13.3 billion dollar valuation, with more than one third of its monetization-focused users already earning revenue.
CoreWeave revenue doubled over the past year to 2.6 billion dollars for the quarter, but cash burn also doubled to 5.7 billion dollars per quarter, and the company reported a 104 billion dollar demand backlog that grew by an additional 25 billion dollars after books closed at end of June. Nebius recorded 454 percent revenue growth to reach 582 million dollars, beat analyst earnings per share forecasts by 83 percent, and saw Blackwell compute auction prices clear at 15 percent above the previous record for Hopper compute. Tencent spent 7.8 billion dollars on AI infrastructure in the past quarter while reporting negative free cash flow, and the broader observation is that China's AI buildout narrative is tracking approximately three to six months behind the US narrative arc, with US market participants not fully accounting for a Chinese capex boom.
Google DeepMind CEO Demis Hassabis and longtime product leader Jeff Dean departed last week. Sergey Brin addressed a town hall after the release of Mythos, telling engineers it is time to play catch up, and has used his implicit power as co-founder to push resource allocation toward areas including recursive self-improvement. Google teams are reported to be shifting focus from Gemini 3.5 Pro to the scaled-up Gemini 4, with the view that releasing a competent Gemini 3.5 Pro at this point would be seen as a failure given how far behind schedule it already is. The Trump administration's model safety testing framework is expected to expand to open models once they reach capabilities equivalent to models like GPT-5.6, with President Trump reportedly insisting the framework remain voluntary on the grounds that formal regulation would help China catch up.
This summary was generated from the episode transcript and can contain mistakes.