PodBrowser
Practical AI

AI incidents, audits, and the limits of benchmarks

Friday, 13 February 2026 · 3 min read · Listen to the episode ↗

This conversation highlights three key topics: the significance of understanding AI incidents to enhance safety practices through learning from past failures, the necessity for third-party audits to ensure trustworthiness in AI systems, and the challenges surrounding the definitions and applications of audits versus benchmarks in AI evaluations. Insights about maintaining data integrity through mandatory reporting and collaborative efforts in AI vulnerability assessment underscore the importance of safety in the evolution of AI technologies.

Sean McGregor, co-founder of the AI Verification and Evaluation Research Institute, emphasizes the importance of understanding AI incidents and their consequences. He advocates for learning from past AI failures to enhance safety practices, drawing parallels to established safety protocols in aviation and food industries. His work includes the AI Incident Database, which has collected over 5,000 human-annotated reports, highlighting the impact of AI incidents on businesses, including stock price fluctuations.

The conversation reveals the lack of a standard definition for safety in AI, with nuanced meanings for terms like incident, accident, and harm event. Sean defines an "incident" as an event where harm has occurred, advocating for a focus on significant incidents that can inform safer AI production. He notes challenges in categorizing incidents and the potential for data overload, while acknowledging the increasing scale of psychological harms on populations.

Data for the incident database primarily comes from journalistic reporting, which is crucial for maintaining integrity. Sean discusses the decline of the journalistic community and its implications for reporting on AI incidents, advocating for mandatory reporting, especially in light of the EU's code of practice for severe incidents. He argues that mandatory reporting yields more insightful data, despite challenges in assigning rates to incidents over time.

The conversation also touches on the security implications of AI models and the necessity of third-party audits to verify safety. Sean points out the difficulties in assessing general-purpose AI systems from companies like OpenAI and Google due to their broad operational contexts. He stresses the need for practical evaluations, such as pilot programs, to ensure AI systems can avoid catastrophic outcomes, acknowledging that third-party evaluations may provide more reliable assessments than first-party ones.

A traffic camera incident illustrates the complexities of AI incidents, raising questions about the application of processes in larger instances and the differences between third-party and internal evaluations. The importance of audits for organizations, particularly mid-sized companies, is emphasized, noting that while audits are often unpopular, they are crucial for building trust and attracting investments.

The discussion distinguishes between audits and benchmarks, explaining that audits verify financial representations, while meta-evaluation assesses the validity of benchmarks. Many benchmarks are designed for research rather than practical application, leading to challenges when organizations attempt to implement them in real-world scenarios. The BBQ benchmark is mentioned as significant in the bias benchmarking community, but its limitations for specific AI applications are acknowledged.

The conversation includes references to the Darknet Diaries podcast and the AI village at Def Con, where participants engaged in challenges to expose vulnerabilities in AI models. The Generative Red Team 2 challenge required systematic analysis to ensure findings were statistically significant, emphasizing the need for collaboration across security, safety, statistics, and machine learning disciplines to enhance AI reliability.

Unexpected modes of failure in AI systems, particularly regarding guard models and foundation models, are highlighted. There is a strong emphasis on understanding the integration and configuration of these models, noting issues such as the hard rejection of non-compliant prompts and the loose configuration of model handoffs, which can lead to exploitation. The unpredictability of interactions between multiple systems underscores the necessity for thorough testing of interfaces.

Participants discussed the value of documenting flaws in AI systems, suggesting the implementation of tools similar to bug bounty systems in traditional computer security. The research value generated from activities like DEF CON was recognized, emphasizing the need for a structured flaw reporting system tailored for machine learning and data-driven systems.

Sean reflected on the future of his work and the broader research ecosystem, expressing optimism about community progress while acknowledging ongoing complexities in AI safety. The conversation questioned the appropriateness of celebrating milestones in AI safety when underlying issues persist, with a collective desire for a future where safety problems are resolved, alongside a recognition that complexities will continue to emerge. There is a call to prioritize the deployment of safe systems to clients who value safety and outcomes, along with a commitment to improving risk assessment measures.

This summary was generated from the episode transcript and can contain mistakes.