How Foundation Models Evolved: A PhD Journey Through AI's Breakthrough Era
Tuesday, 18 November 2025 · 3 min read · Listen to the episode ↗
Omar Khatab explores the evolution of AI foundation models, emphasizing the importance of human-designed systems that prioritize effective communication, retrieval, and training while critiquing the belief that larger models will lead to AGI. He advocates for programmable tools and a structured framework like DSPY to clarify user intent and enhance AI's reasoning capabilities. Khatab also discusses the need for collaboration in AI development and the potential implications for user discernment in the context of evolving model capabilities.
Omar Khatab emphasizes the need for specific solutions in AI, cautioning against over-engineering and advocating for human-designed pipelines that prioritize retrieval, web search, and agent training. He critiques the belief that larger language models will inevitably lead to artificial general intelligence (AGI), suggesting a focus on artificial programmable intelligence instead. Khatab raises important questions about how humans can effectively communicate their needs to scalable models, highlighting the ambiguity of natural language and the rigidity of code. He calls for a new abstraction layer to improve communication between users and AI systems.
Khatab discusses the importance of creating programmable tools for AI that enable reasoning and composition, rather than treating AI as an inscrutable oracle. He reflects on his PhD experience, where he initially resisted the idea that scaling model size and pre-training would solve all problems, arguing for systems that extend beyond mere model capabilities. He acknowledges the rapid advancements from frontier labs and the shift away from the belief that scaling parameters and pre-training data are sufficient.
The conversation touches on understanding user intent in AI, critiquing the alignment approach for its communication shortcomings. Khatab believes that models will improve significantly over time with the right context and instructions, advocating for smarter software systems rather than equating AGI with human or animal intelligence. He stresses the need to identify specific problems for AI to address, such as enhancing medical research and healthcare outcomes.
Khatab critiques the notion of waiting for models to improve, arguing that complex user desires cannot be captured in simple prompts. He identifies three reasons for this complexity: users may lack clarity about their wants, desires can be intricate, and fundamental trade-offs often lead to ambiguity. The discussion raises questions about whether powerful models might lead individuals to become less discerning about their wants, potentially abdicating their preferences to models for recommendations.
The speaker distinguishes between recommendation algorithms and search algorithms, emphasizing that true intelligence involves a nuanced understanding of the world and human interests. They introduce DSPY, a framework designed to enhance prompting for complex applications, highlighting the challenges of prompt engineering and the need for effective communication with models. The development of DSPY over the past three years is discussed, noting that effective prompt engineering requires aligning requests with model capabilities.
Khatab introduces the concept of "signatures" in DSPY, which manage ambiguity in interactions with language models by isolating it into functions. These signatures require users to articulate their intent clearly, combining technical elements with descriptive language. The importance of specifying document types within signatures is highlighted, as different types carry distinct semantic meanings.
The conversation also addresses multi-agent systems and inference time strategies, advocating for expressing intent in its most natural form. Khatab critiques the misuse of model power, asserting that clear communication of needs is more effective than relying solely on models. He explains that DSPY's database of optimization aims to streamline processes and reduce extensive exception lists.
The discussion highlights the need for structured abstraction in AI software engineering, questioning whether natural language can serve as a complete specification. Khatab warns that poorly constructed systems can lead to ineffective programs, stressing that the goal of building optimizers is to help users express their intent more clearly. He introduces declarative programming, which allows users to define desired end states without managing every event, simplifying programming for complex systems.
Khatab discusses the evolution of optimization methods, detailing the transition from early models that struggled with instruction following to more advanced reflective prompt optimization techniques. He emphasizes the importance of signatures in natural language, structured control flow, and data in fully specifying user intent, acknowledging the challenges of mapping complete problems into these components.
The speaker advocates for an open-source structure in AI development, arguing that collaboration among academics and researchers will enhance all programs. A philosophical question is raised regarding whether models will develop independent agency or remain guided by humans, with Khatab suggesting that the necessity for humans to declare intent may diminish as AI evolves. He concludes that both structured systems and flexible interactions are essential, highlighting the distinct roles of software systems and human-like agents.
This summary was generated from the episode transcript and can contain mistakes.