PodBrowser
a16z AI

Evals, Feedback Loops, and the Engineering That Makes AI Work

Tuesday, 17 February 2026 · 3 min read · Listen to the episode ↗

The conversation delves into the engineering essentials of AI, emphasizing evaluation (evals) and feedback loops, crucial for refining models rather than merely employing brute force methods. Ankur Goyal critiques the AI industry's reliance on data and computational power over traditional engineering practices, highlighting concerns around model complexity and integration into enterprises. The discussion also touches on market dynamics and challenges in adopting AI technologies, alongside investor pressures on growth metrics and profitability.

Ankur Goyal emphasizes the distinction between AI's continuous nature and the discrete systems humans typically design, which prioritize predictability and reliability. He critiques the AI industry for often resorting to brute force methods, where teams prioritize the latest models over refining existing ones. Successful AI product companies excel in engineering practices, particularly in evaluations, feedback loops, and testing harnesses. Goyal defines "eval" as applying the scientific method to AI systems, involving hypothesis testing and output observation to measure improvements, stressing the need for both quantitative and qualitative assessments.

The conversation explores the divide between AI specialists and systems product experts, with Goyal identifying himself as a systems product person transitioning into AI. He highlights the tension between AI's focus on optimization and non-determinism versus systems' emphasis on correctness and reliability. The "bitter lesson" concept suggests that while AI models are powerful, they often rely heavily on data and computational power, sometimes at the expense of traditional engineering principles.

The speakers discuss the importance of engineering efforts in data handling, noting that significant work is required to make models usable, especially during training. They express concerns about managing complexity in AI development, particularly when substantial resources are invested in training models. Speaker 1 expresses skepticism about whether a deeper understanding of AI models translates to improved efficiency, advocating for the scientific method and effective frameworks for AI.

Speaker 2 concurs on the importance of problem understanding and notes that AI models disseminate intelligence rapidly once released. They raise questions about the distribution of intelligence from these models, particularly regarding Chinese models, and mention a benchmark comparison of various models, highlighting trade-offs in cost, error rate, and latency. The conversation touches on the low market share of certain models and the trend of choosing cheaper models for cost savings, contributing to self-cannibalization.

The discussion anticipates a shift in model distribution as frontier models slow down, with a focus on innovative releases. There is a noted tendency to overlook open-source models in favor of new frontier models, despite some customers continuing to use older models for their effectiveness. Performance in AI is defined by latency and accuracy, with teams optimizing existing models for specific use cases. Familiarity with the model is emphasized as crucial for achieving optimal performance.

The rapid implementation of AI technologies is crucial for launching new businesses, although enterprise adoption remains in its early stages. There is a strong demand for AI systems, with confidence expressed in the performance of certain models on relevant benchmarks. However, integrating AI into enterprise settings presents challenges, and while organizations may struggle, individuals are increasingly utilizing tools like ChatGPT and Gemini.

Market dynamics suggest a competitive landscape, resembling a two-horse race, with different saturation levels for individual users and enterprises. Future AI capabilities could extend to everyday tasks, but current automation solutions are lagging behind text prediction models. Technical challenges in AI development are highlighted, with efforts to improve efficiency through batching LLM calls and addressing coding problems across multiple programming languages.

Investor perspectives reveal that founders are influenced by expectations to focus on growth metrics rather than profitability, complicating the path for startups. Ankur reflects on strategic frameworks for managing growth and margins in cloud services, emphasizing the importance of aligning pricing with customer usage. Concerns about potential fraud in token reselling are raised, with the speaker noting that their product's higher entry price helps mitigate abuse.

The discussion highlights the efficiency of SQL over bash for data organization and access, recognizing SQL as more accurate and faster. Two perspectives on technology usage emerge: one that favors brute force computing and another that emphasizes a solid understanding of computer science fundamentals. The potential for a golden age in computer science is noted, particularly regarding tools that enhance reliability and correctness. The conversation wraps up positively, with both the host and guest expressing appreciation for the discussion.

This summary was generated from the episode transcript and can contain mistakes.