PodBrowser
The AI Podcast

Lowering the Cost of Intelligence With NVIDIA's Ian Buck - Ep. 284

Monday, 29 December 2025 · 3 min read · Listen to the episode ↗

The discussion features Ian Buck of Nvidia delving into the Mixture of Experts (MOE) architecture, which enhances AI model performance while lowering operational costs through selective neuron activation. Buck compares traditional models with MOEs, highlighting their efficiency in delivering intelligent responses and emphasizing the impact of GPU advancements on AI capabilities. Additionally, the conversation addresses the significance of tokenomics in reducing costs amid complex AI systems, pointing to IoT advancements as vital for future applications.

Noah Kravitz interviews Ian Buck from Nvidia, focusing on the Mixture of Experts (MOE) architecture and its significant impact on AI model performance and cost. Ian explains that MOE allows models to activate only the necessary neurons for specific tasks, contrasting with traditional neural networks that require all neurons to be active, which can slow performance and increase costs.

He compares models like Llama, which has 405 billion parameters but activates all of them, to OpenAI's GPT model, which activates only a fraction of its 120 billion parameters yet achieves a higher intelligence score. This illustrates how MOE can enhance intelligence while reducing operational costs, despite its complexity in implementation.

The discussion explores how MOEs function, with experts emerging from data rather than being hard-coded. A router directs inquiries to the appropriate experts, allowing for parallel processing. Ian uses the analogy of training multiple domain experts to illustrate the efficiency of this approach. Historical context reveals that while the concept of using multiple models for improved accuracy is not new, the application of MOEs in AI has gained traction recently, particularly marked by the "Deep Seek" moment, which showcased a competitive MOE model achieving high intelligence scores and cost-effectiveness.

Ian notes that models designed for intelligent responses benefit from being MOEs, as they can encode extensive knowledge efficiently. Larger models can activate only a small percentage of neurons, allowing for more experts while managing costs, although communication among them can be complex. The conversation also touches on "tokenomics," emphasizing the importance of reducing costs while maintaining or improving intelligence in AI systems.

Ian explains that while complex AI systems may have higher initial training costs, they can ultimately lower total costs through improved efficiency. The relationship between AI hardware advancements and model capabilities is highlighted, showing how progress in GPU technology has enabled the development of more sophisticated AI models. The evolution of GPU technology is discussed, emphasizing the transition from basic PCIe cards to advanced GPUs that enhance floating point calculations and memory efficiency.

The concept of Total Cost of Ownership (TCO) is introduced, focusing on the cost of intelligence per dollar and the goal of reducing operational costs over time. Ian discusses Nvidia's commitment to incorporating new technologies in each architecture generation, particularly the introduction of HPM memory, which, despite its higher cost, provides significant performance benefits. The DeepSeek R1 GPU system, utilizing the Hopper H200 with eight GPUs connected via MeLink, is highlighted for its ability to function as a single large GPU, enhancing efficiency.

Ian addresses the hidden costs related to communication among experts in MOE, emphasizing the need for rapid communication to avoid idle time. The introduction of MVLink is presented as a solution to communication bottlenecks, allowing GPUs to communicate at full speed. The Kimi K2 model, a trillion-parameter model, operates efficiently with only 32 billion parameters for responses, demonstrating cost efficiency through MVLink connectivity.

Nvidia's collaboration with AI companies to build data centers and maximize GPU utility is emphasized, alongside ongoing software development efforts to enhance model performance. A recent success story illustrates the effectiveness of NVFP4 techniques, achieving a 2x performance increase in just two weeks. The complexity of managing 72 GPUs with a large team of experts is discussed, highlighting the critical need for extreme code design that integrates hardware, model builders, and the software stack.

Concerns are raised about the potential risks of over-focusing on MOE and its future relevance. The conversation stresses the necessity for models to evolve, incorporating reasoning techniques to unlock new opportunities across various industries. The broader impact of AI is discussed, with the supercomputing community increasingly adopting AI for applications in fields such as physics and weather simulations, as well as in biology and drug discovery.

Looking ahead, there is excitement about the ongoing development of MOEs and the goal of reducing token costs. While advancements may lead to more complex and costly technology, they are expected to enhance overall capability and intelligence. The conversation concludes with a focus on the importance of building and supporting customers in this evolving landscape.

This summary was generated from the episode transcript and can contain mistakes.