PodBrowser
Latent Space

Information Theory for Language Models: Jack Morris

Wednesday, 2 July 2025 · 2 min read · Listen to the episode ↗

Jack Morris discusses his experiences in machine learning, focusing on language models and their rapid evolution since 2019. He highlights the importance of agility in adapting to advancements like GPT-3, emphasizing the need for improved training in high-performance computing (HPC) for graduate students. Central to his work is the concept of embedding and information theory, particularly regarding extractable information in pre-trained models, and the potential for new methodologies to enhance model alignment and efficiency.

Jack Morris, a PhD student at Cornell Tech, shares his journey in machine learning, particularly in language applications, beginning his research in 2019. Influenced by models like AlphaGo, BERT, and GPT-2, he notes the significant changes in the AI landscape since starting grad school in 2021, particularly with the release of GPT-3 and ChatGPT, which transformed public interest in the field. He discusses the challenges graduate students face in adapting to rapid advancements in AI and emphasizes the importance of agility in research, advising students to quickly re-implement new ideas.

Morris highlights the shift in model scale, where larger models demonstrate superior capabilities, and the need for academia to catch up with these advancements. He identifies a gap in training for high-performance computing (HPC) among grad students, who often rely on online resources for GPU training. Mastering GPU usage is deemed essential for enhancing researchers' independence and employability. The introduction of Mojo, a new programming language aimed at improving GPU programming efficiency, is also discussed.

Morris emphasizes the significance of his work on contextual documents and embedding models, referencing a post titled "A New Type of Information Theory." He explains concepts from information theory, such as measuring information with computational power constraints and the idea of "extractable information," suggesting that pre-trained models possess more extractable information than randomly initialized ones. The complexity of model weights and activation vectors is highlighted, along with the challenges in understanding them.

The conversation touches on the relevance of comparing language model storage to Wikipedia, emphasizing the importance of understanding "usable information" in language models. Morris shares his research on recovering data from compromised vector databases, achieving significant recovery rates, and discusses debiasing embeddings. He reflects on his research breakthroughs and the invigorating nature of the research process, expressing gratitude for his graduate school experience.

Morris introduces the "Platonic Representation Hypothesis," suggesting that models trained on similar data converge to learn comparable concepts. He aims to combine this hypothesis with his vector-to-text methods to enhance model alignment and embedding inversion. The challenges of embedding large texts into smaller vectors and the implications of lossy compression are also discussed, along with the concept of a "cognitive core" for efficient on-device use.

The conversation explores the need for a challenge or leaderboard to assess model performance, emphasizing that the best memorization model may not equate to the best generalization model. Participants stress the importance of asking new questions in research to foster innovation. They discuss tools for creating visualizations and the significance of professionalism in published work.

Morris and his co-speaker reflect on the importance of datasets in AI and the cycles of innovation in the field, identifying four significant paradigm shifts in language models. They speculate on future innovations arising from new data sources and acknowledge the ongoing debate about the importance of data and compute efficiency. The conversation concludes with an invitation for speculation on future advancements in AI and language models.

This summary was generated from the episode transcript and can contain mistakes.