Long Horizon Agents, State of MCPs, Meta's AI Glasses & Geoffrey Hinton is a LOVE RAT - EP99.17
Thursday, 18 September 2025 · 4 min read · Listen to the episode ↗
The podcast discusses **Model Creation Processes (MCPs)**, including innovations like the **Audiobook Maker** and **Cdream 4**, while addressing model quality issues. It highlights the evolution of **Long Horizon Agents** in AI aiming for improved autonomy and error correction, emphasizing the need for supervision. Additionally, **Meta's AI glasses** are explored for their potential and challenges in providing contextual awareness, alongside commentary on Geoffrey Hinton's personal revelations, blending AI discourse with human interest.
Chris discusses new Model Creation Processes (MCPs), particularly **Cdream 4**, which creates stunning images, and the **Audiobook Maker MCP**, which converts stories into audiobooks. He raises concerns about the quality of models, questioning whether providers intentionally degrade quality for cost savings. Anthropic, now known as Claude, acknowledges issues with their model's effectiveness, including a routing error that affected up to 16% of requests.
Recent research titled "The Illusion of Diminishing Returns" suggests that execution is the main bottleneck in large language models (LLMs). Errors during task execution can worsen subsequent outputs, but correct reprompting can lead to successful task completion. The hosts reflect on their experiences, noting that guiding AI conversations can improve responses while expressing concerns about AI systems compounding errors without human intervention.
The conversation includes an AI village experiment where AIs pursue long-term goals, emphasizing the need for advancements in autonomous agents and the importance of a supervisory agent to identify when an AI goes off track. The capabilities of GPT-5 are noted, particularly its ability to execute over a thousand steps correctly, although concerns about high latency are raised. The speakers propose a voting system among expert agents to determine the best course of action and discuss the balance between intervention strategies.
The need for clear job descriptions and defined tasks for AI agents is emphasized, ensuring security and permissions are established. The limitations of complex models with numerous tools are highlighted, as they may struggle to select the appropriate tools for specific tasks. A trained AI system is deemed more valuable, capable of following established steps and learning from past experiences.
The gradual development of AI autonomy is explored, with one speaker underscoring the need to equip AI systems with effective tools. They introduce the "dry run" concept, where an AI agent summarizes its intended actions before proceeding, emphasizing the importance of human supervision for timely feedback. The current technological limitations in achieving full autonomy are acknowledged, alongside the necessity for context, interaction, and reliable performance from AI systems.
Training AI for specific tasks, such as solving support tickets, is discussed, with the goal of achieving an 80% success rate before granting more autonomy. The potential for AI to manage busy work in business processes is seen as a significant advantage. The efficiency of AI in handling support tickets is noted, as it can quickly gather relevant information from multiple systems, saving time compared to manual processes.
The speaker discusses the internal tool Sim MCP, which facilitates custom plan creation and user discovery. They emphasize the need for exposing business context and data analysis, highlighting the complexity of data spread across various systems. The integration of LLMs with MCPs and internal data is suggested as a way to streamline data analysis. Challenges in synthesizing data from multiple systems and APIs are noted, along with a decline in trust in tools like ChatGPT for data analysis.
The current state of MCPs is described as disorganized, necessitating a comprehensive registry to improve consistency and usability. The speaker expresses skepticism about drag-and-drop tools for building reliable autonomous use cases, advocating for agents that function as support workers with clear instructions. Market challenges are noted, including the hype surrounding MCPs that can lead to disappointing user experiences.
The podcast discusses the potential of AI tools for developers, emphasizing the creation of a "dream tool" to enhance problem-solving. Dario from Anthropic predicts that a significant portion of code will be generated by LLMs, although there is skepticism about the timeline. Concerns about AI's impact on jobs are addressed, with the belief that developers and support workers will adapt rather than be replaced.
The discussion transitions to Meta's Ray-Ban display, which features cameras, audio, and a built-in assistant. Users can ask questions and receive information, but the primary function is as headphones. New features include a small screen for directions and discreet text replies. Observations indicate that users may appear distracted while using the glasses, raising concerns about maintaining focus.
The podcast critiques the current use of AI in the glasses, suggesting missed opportunities for passive context gathering. Potential applications include systems that infer environmental context. In business and sports, the glasses could provide valuable applications, with one speaker expressing interest in features like video calls and integration with fitness apps.
The conversation explores the potential of Meta's AI glasses, with varying opinions on whether they signify the future of computing or are merely a novelty. Jeffrey Hinton, known as the "AI Godfather," is discussed in light of recent media attention surrounding his breakup, where his ex-partner utilized ChatGPT to analyze his behavior. Hinton's candid admission of being a "love rat" reveals a personal side to his public persona, leading to playful commentary among the speakers.
This summary was generated from the episode transcript and can contain mistakes.