Mira Murati's 975B Open Model, Ramin Hasani on Post-Transformer AI, and Demis' AI FINRA | EP #271
Friday, 17 July 2026 · 4 min read · Listen to the episode ↗
This episode covers Mira Murati's Thinking Machines Labs releasing Inkling, a 975 billion parameter mixture-of-experts open weight model firing 41 billion parameters at inference and trained on 45 trillion tokens across text, image, audio, and video, with a business strategy centered on reinforcement fine tuning for enterprise customization rather than leaderboard dominance.
Mira Murati's Thinking Machines Labs released Inkling, a mixture-of-experts open weight foundation model with 975 billion total parameters that fires 41 billion at inference time and was trained on 45 trillion tokens spanning text, image, audio, and video. Murati's stated strategy is customization over leaderboard dominance, and the company's own blog does not claim Inkling is the best model in the world. On released evals, Inkling scores stronger than Neymetron but weaker than GLM 5.2 and weaker than closed-weight Western frontier models. Reuters framed the release as a Western alternative to Chinese open weight models like DeepSeek, and the hosts noted the well-worn industry path of releasing open weight models to generate attention and data before going closed to launch a profitable API business.
The Thinking Machines business model rests on reinforcement fine tuning as the dominant paradigm for model customization. Unlike supervised fine tuning and LoRA-style adapters, which the speakers describe as producing style transfer at best without raising capabilities, reinforcement fine tuning is characterized as the first fine tuning method that demonstrably increases capability levels. Ramin Hasani argued the market for model customization is severely undersupplied, that OpenAI's fine tuning API never gained traction and has since been wound down, and that fine tuning at scale could generate one to two orders of magnitude more tokens than base model inference alone. A significant caveat is that reinforcement fine tuning is probably not the permanent long-term solution, and a sufficiently capable generalist model approaching artificial superintelligence may eventually not benefit from further fine tuning on internal data sets, which would undermine the core business premise.
Demis Hassabis published an essay calling for a US-led frontier AI standards body modeled on FINRA that would test frontier models before release, reportedly wanting the body operational before the end of the current year. Sam Altman published a parallel op-ed in the Financial Times proposing a US-led international forum for AI standards and risk assessment. Alex argued the FINRA analogy has surface plausibility but that the differences outweigh the similarities, and characterized frontier lab CEO calls for regulation as smelling like regulatory capture and an attempted cartel formation among incumbent labs that could box out open weight, open source, university-driven, and non-incumbent frontier models. Salim rated the question of whether the calls reflect genuine safety concern versus regulatory capture as roughly 50-50. Bill added that no one at a fast-moving AI lab would carve out years to participate in such a body the way finance executives do in FINRA, and that AI can help regulate itself in a way that has no equivalent in finance.
Alex identified two distinct regulatory frames: regulating model capabilities at construction time versus regulating the actions models take in deployment. He argued that capability-level regulation is tantamount to thought policing the models and that the only plausible enforcement mechanism would require a globally agreed pandemic-style threat detection scheme requiring alignment between the US and China at minimum. The White House is reportedly weighing a framework that would clear US models for release as long as they stay at or below the level of China's best open weight model, with Chinese open weight models currently trailing US models by an average of seven months. Alex argued this creates a perverse incentive to let China advance to ever greater capability levels so Western labs can escape regulation, and that freezing capability benchmarks could distort model development by over-incentivizing certain capabilities and under-incentivizing others.
Wiko AI, founded by University College London researchers, published experimental evidence claiming the first recursive self-improvement system, in which an outer AI agent rewrites the code and research strategy for an inner AI agent. Wiko claims eight days of machine self-improvement beat two years of expert human effort and rates their system at level one on a zero-to-three scale analogous to SAE autonomy levels. Hasani was less enthusiastic, arguing that no weight changes occur in the neural networks during the Wiko pipeline, meaning core model competencies remain fixed, and that true recursive self-improvement requires AI systems that can retune their own weights. He outlined a hierarchy from prompt engineering at the shallowest level, to fine tuning a smaller version of itself, to pre-training the next generation of itself as the deepest level, and noted that Andrei Karpathy joined Anthropic to work on pre-training automation. Hasani predicted unbelievably capable models surpassing human understanding will likely appear within two years, contingent on continued compute growth and the absence of chip or memory shortages.
Ramin Hasani's Liquid AI builds on liquid neural networks derived from studying the 302-neuron nervous system of C. elegans, using recurrent neural networks and continuous time processes rather than attention mechanisms. Liquid AI claims its models can deliver intelligence at the level of models 10 to 1000 times larger and can run on CPUs outside data centers. The company deployed a sub-one-gigabyte model in Mercedes-Benz vehicles running on chips costing as little as 60 dollars, operating fully offline, with an over-the-air update rolling out this year to 2022-onward vehicles. Liquid AI has also served one billion requests inside Shopify's framework across hundreds of millions of users and 10 billion products over six months.
GPT-5.6 set a new all-time high on OpenAI's HealthBench Professional benchmark of 525 real clinical tasks, outperforming specialty-matched physicians with unlimited web access in blind evaluation drawing on roughly 20,000 physician judgments.
This summary was generated from the episode transcript and can contain mistakes.