What's Missing Between LLMs and AGI - Vishal Misra & Martin Casado
Tuesday, 17 March 2026 · 3 min read · Listen to the episode ↗
The discussion highlights the limitations of large language models (LLMs) in achieving artificial general intelligence (AGI), emphasizing the need for continual learning and a shift from correlation-based learning to causation. Misra illustrates how LLMs function, using examples like retrieving database queries and Bayesian updating, while stressing their fixed nature post-training. The conversation also explores the theoretical aspects of building causal models to advance understanding beyond mere correlations, crucial for developing AGI and innovative frameworks in machine learning.
Vishal Misra discusses the capabilities and limitations of large language models (LLMs) in relation to artificial general intelligence (AGI), emphasizing that LLMs lack consciousness and inner monologue despite their impressive outputs. He proposes a thought experiment where an LLM trained on pre-1916 physics deriving the theory of relativity would indicate AGI. Misra identifies two necessary advancements for achieving AGI: the ability to continue learning post-training and a shift from correlation-based learning to understanding causation.
Misra shares his early experiences with GPT-3, particularly in solving a problem related to querying a cricket database. He describes implementing retrieval augmented generation (RAG) in 2020, which allowed GPT-3 to translate natural language for database queries. The conversation highlights the distribution of next tokens in LLMs, illustrating how choices can significantly alter context. Misra explains Bayesian updating, where prior knowledge is adjusted based on new evidence, and discusses the sparse nature of the matrix due to many nonsensical combinations. In-context learning is also addressed, showcasing how LLMs can learn in real time by being shown examples they haven't encountered before.
Misra discusses using GPT-3 to create a front-end for the Stats Guru database by designing a domain-specific language (DSL) to convert natural language queries into this format. He acknowledges the ongoing debate about characterizing LLMs as Bayesian and expresses a desire to clarify this aspect with support from the field. The conversation introduces the concept of a "Bayesian wind tunnel," likening it to aerospace testing environments, revealing that transformers excelled in matching Bayesian posteriors.
The discussion shifts to the comparison between human cognition and machine learning. Both humans and models engage in Bayesian updating, but human brains exhibit lifelong plasticity, allowing for adaptive learning. In contrast, LLMs, once trained, have fixed weights and do not retain information from past interactions. A key distinction is made between human decision-making and deep learning models, highlighting that deep learning focuses on correlations rather than causation.
The limitations of current deep learning architectures are discussed, particularly in relation to Shannon entropy and Kolmogorov complexity. The speaker asserts that deep learning remains within the realm of Shannon entropy, lacking a transition to causal understanding. They stress that achieving AGI requires addressing two main challenges: implementing plasticity through continual learning and shifting from correlation to causation. The conversation concludes with a reference to the "Einstein test" for AGI, illustrating that a model focused solely on correlations would not have led to revolutionary scientific insights like Einstein's theory of relativity.
The discussion also contrasts Kolmogorov and Shannon entropy, emphasizing the need for new representations to effectively describe complex data. While LLMs can simplify complex phenomena, they lack the ability to create new frameworks or manifolds, which is essential for deeper understanding. AGI is defined through various lenses, including the ability to pass the Turing test and perform economically useful tasks. Current LLMs often require human intervention, but advancements may enable them to autonomously handle well-defined coding tasks.
The conversation touches on the concept of chromograph complexity, which remains largely theoretical, with no practical algorithms available for finding the shortest program. Building causal models is identified as crucial for transitioning from correlation to causation, which is vital for advancing intelligence. Recent research indicates a shift in perspective towards these models, with a focus on enhancing plasticity and developing causal understanding. Judea Pearl's causal hierarchy and do calculus are highlighted as valuable frameworks for exploring causality in relation to simulation.
This summary was generated from the episode transcript and can contain mistakes.