Better Data is All You Need — Ari Morcos, Datology
Friday, 29 August 2025 · 2 min read · Listen to the episode ↗
In the discussion with Ari Morkos of Datology, the top insights focus on the critical role of data quality in AI and machine learning, encapsulated in the notion that "garbage in, garbage out." Morkos emphasizes the need for improved data curation methods, particularly in light of challenges from self-supervised learning and synthetic data. He also discusses the economic factors related to open-source datasets and the importance of domain expertise for enhancing model performance.
Ari Morkos, CEO and co-founder of Datology, discusses the company's mission to enhance data quality in machine learning, emphasizing that "models are what they eat." He advocates for automation in data management to improve model training speed and performance, aiming to make data curation accessible to non-experts. Morkos shares his transition from neuroscience to machine learning, influenced by AI advancements, and highlights the challenges of optimizing models based on correlates rather than causal variables.
He critiques the misconception that all data is equal and stresses the importance of data quality, echoing "garbage in, garbage out." Morkos notes the shift in the field with self-supervised learning, which has increased data availability but also led to issues with low-quality data. He emphasizes the need for improved data curation methods and the importance of identifying the most informative data points for model training.
The conversation touches on the variability in data needs, illustrating that different concepts require different amounts of data. Morkos discusses the complexities of data curation, including the challenges of human-guided efforts and the need for a holistic view of the training process that integrates pre-training, mid-training, and post-training phases.
He raises concerns about the economics of open-source datasets and the balance between open-source contributions and proprietary enhancements. Morkos identifies scientific know-how, engineering infrastructure, and brand recognition as potential sources of competitive advantage for Datology, which aims to train models faster, better, and smaller.
The discussion highlights the significance of synthetic data in pre-training and the need for diversity in data quality. Morkos critiques the reliance on textbooks for quality data, noting that narrow distributions can hinder generalization. He emphasizes the importance of developing effective curricula and the need for data analysis in research culture.
Morkos reflects on the challenges of acquiring domain expert data and the necessity for curation to adapt to different datasets. He invites data enthusiasts to join Datology, emphasizing the critical role of data in AI and the potential for significant performance improvements through better data quality. The conversation concludes with recognition of key figures in the field and the importance of investing in data and talent as essential compute multipliers.
This summary was generated from the episode transcript and can contain mistakes.