PodBrowser
Latent Space

[AIEWF Preview] Gemini in 2025 and Realtime Voice AI

Monday, 2 June 2025 · 2 min read · Listen to the episode ↗

The podcast features discussions on significant advancements in AI, particularly the Gemini model and real-time voice technology. Key updates include Google's innovative features like multilingual audio output and text-to-speech developments. The challenges of productionizing generative UI for dynamic website building and the complexities of real-time voice agents are also addressed, alongside aspirations for Gemini 5.0 to enhance language support. Cryptocurrency and blockchain technology aren't directly mentioned.

Sam Ritarnick introduces guests Swix and Logan on the TwiMul AI Podcast, where they discuss significant advancements in AI, particularly focusing on Gemini and real-time voice technology. Logan highlights key updates from Google I.O., including "thinking budgets" for the 2.5 Pro model and the integration of thought summaries into Cursor. He expresses enthusiasm for a new native audio output feature that supports multiple languages and high-quality voices, along with a tool called URL context for developers to access detailed web information.

The conversation delves into Gemini Diffusion's potential for generative UI, which enables dynamic website building based on user interactions, though challenges remain in productionizing it due to the need for high-quality models and token generation speed limitations. The significance of audio and video in AI applications is underscored, particularly in transcription for the live API, which presents challenges such as session length limitations and the need for higher developer commitment compared to simpler models.

Logan discusses the introduction of new high-performing text-to-speech models and Google's goal with the Gemini model to create a unified system that enhances performance through multimodal capabilities, including video understanding. The podcast also covers the integration of autoregressive and diffusion models in image generation, with developers working to incorporate these into Gemini while utilizing existing models.

The design of Gemini's APIs is a focal point, emphasizing a shift to a component-based architecture for improved quality and low latency. Challenges in voice technology, such as balancing latency, cost, and output quality, are acknowledged, along with advancements in voice activity detection models that allow for customization. The complexities of building real-time voice agents are discussed, with collaboration with DeepMind aimed at addressing issues like turn detection and context management.

The podcast highlights a growing interest in networking protocols among developers due to advancements in voice AI and video, with expectations for real-time responses targeting around 500 milliseconds. An experimental feature called "proactive audio" is introduced, designed to filter out irrelevant audio during conversations. Challenges in speaker identification and diarization are addressed, along with the introduction of a new "dialogue" model for improved speaker recognition. The conversation concludes with discussions on future developments for Gemini 5.0, including aspirations for broader language support and plans for further discussions at the World's Fair.

This summary was generated from the episode transcript and can contain mistakes.