Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
Wednesday, 9 September 2026 · 2 min read · Listen to the episode ↗
In this episode, Rayan Krishnan advocates for a new approach to AI model evaluation, emphasizing the need for independent third-party benchmarks to counteract biases in self-assessment. He introduces the recursive self-improvement index (RSI) as a tool for comparing models and highlights the importance of integrating evaluations into enterprise workflows. The discussion also touches on the growing investment in AI and the necessity of involving diverse stakeholders in establishing effective standards, while addressing the risks associated with rapid technological advancements.
Rayan Krishnan argues for a new methodology in AI model evaluation, emphasizing the need for independent third-party companies to create high-quality benchmarks. He asserts that labs cannot effectively evaluate their own models due to inherent biases in reporting capabilities, which can distort public perceptions of AI performance.
Krishnan highlights the rapid evolution of AI models and the corresponding need for updated testing methods, noting that existing benchmarks can quickly become obsolete. He introduces the recursive self-improvement index (RSI) benchmark as a tool for making direct comparisons across different models, underscoring the importance of integrating clear evaluations into enterprise workflows.
The discussion includes a comparison of current measurement challenges in AI to the MPAA's film classification system, illustrating the complexities of establishing effective standards. Krishnan cites a Fortune 10 company's significant increase in its AI budget, suggesting a trend where AI investments may soon exceed traditional salary costs.
He also mentions the VALS Smith tool, which allows companies to create internal coding benchmarks from their GitHub repositories, predicting a market shift towards such tools for AI model evaluation. Krishnan stresses the necessity of involving diverse stakeholders in defining AI standards and the role of third-party evaluators in developing technology that tests these standards.
The episode addresses the tension between rapid technological advancement and the need for public interest safeguards. Krishnan points out that government agencies often receive warnings from labs about potential AI risks and advocates for regular briefings to keep policymakers informed about AI capabilities and associated dangers.
Krishnan discusses the increasing investment in sovereign AI, despite the challenges of building data centers and replicating data processes. He emphasizes the importance of a shared language for evaluations to effectively assess AI risks and capabilities, drawing parallels to verification challenges in nuclear arms control.
He warns that recursive self-improvement in AI could lead to significant disparities in capabilities among countries and companies, each with varying focuses on biosecurity and cybersecurity. Vals.ai is mentioned as an organization dedicated to creating benchmarks that accurately reflect the forefront of AI capabilities and risks.
Krishnan concludes by asserting that evaluations must address new risks at the infrastructure level, rather than just focusing on code vulnerabilities. He believes that the most valuable evaluation businesses will prioritize high-quality assessments that genuinely reflect AI's potential and risks, rather than merely supporting intelligence development.
This summary was generated from the episode transcript and can contain mistakes.