gpt-realtime, nano banana & workspace computer v2 | EP99.15-realtime
Thursday, 28 August 2025 · 4 min read · Listen to the episode ↗
The discussion highlights the recent launch of OpenAI's GPT real-time, which introduces optimized voice assistant features and a notable price reduction, raising concerns about economic viability for businesses. The Gemini 2.5 Image Pro, previously named "Nano Banana," showcases advanced image manipulation capabilities, serving multiple sectors. Additionally, the new Workspace Computer V2 and SimLink app enable remote task management and flexibility, promising enhanced productivity through cloud-based solutions despite challenges with AI model quality and costs.
Chris highlights the chaotic week by wearing a yellow shirt and announces the release of Gemini 2.5 Image Pro and Banana Nano Max. OpenAI's unexpected announcement of GPT real-time introduces new features, including seamless phone purchases via voice assistants, although the examples seem tailored to affluent Californians. The previous GPT 4.0 real-time voice model is critiqued for being cost-prohibitive, raising concerns about the economic viability of voice assistant services for many businesses.
The new real-time API features two voices that are 30% cheaper than the original model, support for remote MCP servers, image inputs, and voice calling through the SIP protocol. There is a potential impact on entry-level call center jobs due to enhanced voice agent experiences, alongside the introduction of emotion detection in voice interactions. A demo from Zillow showcases property search capabilities using the new voice technology, but skepticism arises regarding the practicality of AI features in real-life applications, particularly in phone carrier apps.
Concerns are raised about the quality of responses from voice models, including issues with hallucination and brevity, while improvements in AI's ability to handle tasks like reciting terms and conditions are noted. Speaker 1 questions the effectiveness of Zillow's new features and critiques the notion of making call centers friendlier, advocating for a delegation model that provides concise updates. Speaker 2 suggests testing the new voice technology, leading to a demonstration where the assistant showcases its multilingual capabilities and humorously engages with requests.
The conversation highlights advancements in AI, particularly with GPT real-time, which now allows for asynchronous tool calling and multitasking. The Mike Corp representative emphasizes the model's ability to handle multiple tasks simultaneously, delegating to various assistants and incorporating summarized results into its context without waiting for user input.
There is a discussion about the pricing of GPT real-time, which has seen a 20% reduction, now priced at $32 per million tokens. Concerns are raised about the potential misleading nature of this pricing, especially considering the input/output ratios typically seen in model usage. The speaker suggests that for cost efficiency, an "arm's length" approach should be taken when integrating tools, allowing the model to summarize rather than process large amounts of data directly.
Skepticism is expressed regarding the quality of lesser voice models, with a preference for higher-quality options like Gemini Pro or GPT-5 for productivity. The speaker notes significant trade-offs in quality with current models, indicating reluctance to adopt inferior solutions in business contexts. The conversation also touches on the challenges and costs associated with using AI tools, particularly Firecrawl, which can lead to unexpectedly high bills due to extensive data processing.
Understanding the value of AI is crucial, as many users struggle to grasp the relationship between costs and the value received. The speakers primarily favor GPT-5 for its intelligence and problem-solving capabilities, finding it superior to other models like Gemini 2.5. While GPT-5 is noted for its unique style and data interpretation skills, Gemini 2.5 Pro is preferred for context use.
The introduction of the Gemini 2.5 Flash Image, initially called "Nano Banana," sparked debate over its naming, with some preferring the original. One speaker shares positive experiences with the model's ability to combine images and follow instructions, although limitations exist in handling text and specific modifications. The model excels in creating clear cutout backgrounds and merging images effectively, with potential applications in various fields, including gaming, e-commerce, architecture, and military uses.
The concept of a Workspace Computer is revisited, emphasizing its potential as a cloud-based solution for agents to perform tasks without disrupting user workflows. The team initially faced significant costs while attempting to provide dedicated cloud machines for maintaining agent persistence. This led to a discussion about the disparity between technological possibilities and market realities.
OpenAI introduced the original operator, which evolved into the agent capability on ChatGPT, but automated browsing from cloud server IPs often encounters issues like website blocks and CAPTCHAs. To address affordability and accessibility, the team developed Workspace Computer V2 and introduced SimLink, an app that allows users to create a cloud computer from any device, enabling remote operation from phones or the SimTheory platform.
SimLink facilitates linking multiple computers to a single account, enabling users to delegate tasks to various assistants. The flexibility of SimLink in large organizations allows for short-term leases or individual allocations of machines. The app has advanced capabilities, including reading and writing files and automating repetitive tasks, with users able to train assistants for reliable performance.
Affordable mini PCs can run SimLink headless, serving as virtual cloud PCs. This opens up possibilities for users to create a "swarm" of workspace computers for diverse tasks, with practical applications anticipated in upcoming demos. The conversation shifted to advancements in GPT real-time and Gemini 2.5 image Pro Plus, with excitement expressed about their potential to enhance human-computer interactions.
The conversation centers on local Multi-Channel Processors (MCPs) and their role in enhancing the functionality of workspace computers. Participants express excitement about how MCPs can improve efficiency and user experience, moving beyond traditional mouse and keyboard interactions. The impact of MCPs is recognized as extending beyond just conversational benefits, indicating a broader influence on productivity.
This summary was generated from the episode transcript and can contain mistakes.