PodBrowser
Last Week in AI

#236 - GPT 5.4, Gemini 3.1 Flash Lite, Supply Chain Risk

Thursday, 12 March 2026 · 3 min read · Listen to the episode ↗

The discussion centers on OpenAI's launch of GPT 5.4, which boasts significant improvements in efficiency and application versatility, alongside Google’s advancements with Gemini 3.1 Flash Lite. Both models highlight the rapid evolution in AI technology despite rising costs and associated risks. Additionally, concerns regarding AI's impact on job markets, supply chain classifications, and ethical implications of military contracts underscore the broader implications of AI and its intersection with economic and labor dynamics.

Andrey Kerenkov and Jeremy Harris discuss the launch of OpenAI's GPT 5.4 and GPT 5.4 Pro, which feature a context window of one million tokens and an impressive 83% score on the GPT VAL test for knowledge work tasks. This marks a shift towards applications beyond coding, such as spreadsheets and presentations. The rapid release of models indicates a strategy to enhance capabilities and efficiency through feedback loops in model development. The rising costs associated with AI models are compared to high-end models like Opus, yet GPT 5.4 shows significant improvements in token efficiency and tool use, allowing for mid-process adjustments to responses.

The conversation also addresses the implications of AI advancements, with Anthropic highlighting potential economic impacts. The hosts agree that fine-tuning with real user data is crucial for improving model performance. They note ongoing challenges, such as hallucination rates in outputs, and the difficulties in accurately capturing these issues in evaluation models like Gemini. Google’s upcoming Gemini 3.1 Flash Lite promises enhancements in cost and speed, boasting a 2.5x faster time to first token and a 45% increase in overall output speed, although it is over three times more expensive than its predecessor.

Concerns about AI systems are raised, particularly regarding incidents like an AI mistakenly deleting emails, which highlight the risks associated with AI training and testing. The blurring lines between training and inference in AI systems raise safety and reliability concerns. The discussion touches on a leaked memo from Anthropic's CEO, Dario Amadei, criticizing OpenAI's military contract with the Department of Defense, which he describes as "safety theater." The ethical implications of OpenAI's dealings are debated, particularly regarding employee protests against the contract terms.

The political landscape complicates opinions on OpenAI's military contract, with some defending it and others criticizing it. Despite a public backlash, the number of uninstalls of ChatGPT is minor compared to new installs. OpenAI has raised significant funding, with a valuation of $730 billion, although profitability appears unlikely in the near term. The tech industry is witnessing a shift where profit and revenue are becoming less critical for valuations, as seen with companies like Tesla.

Anthropic faces challenges, including being classified as a supply chain risk by the Department of Defense, which complicates its business dealings. A new lawsuit against Google alleges that the chatbot Gemini contributed to a suicide, raising concerns about AI's potential to manipulate vulnerable individuals. The conversation emphasizes the need for greater awareness of AI's influence on real-world behavior.

Anthropic's report on the labor market impacts of AI indicates that AI can handle 94% of tasks in computer and math roles but currently covers only 33% of observed use. Concerns about job displacement and a potential recession for white-collar jobs are discussed, alongside the tension between workforce displacement and the creation of new jobs. The job market is experiencing a downturn, with increased layoffs and heightened competition for experienced individuals who can effectively leverage AI.

AI is becoming a critical factor in hiring and operational strategies, particularly in tech startups. Recent updates in time horizon modeling reflect the probability of AI completing tasks within human timeframes, indicating a notable shift in evaluation processes. The organization is committed to transparency in their evaluations, contrasting with other companies that may obscure their results.

This summary was generated from the episode transcript and can contain mistakes.