The Agentic Gap: What Enterprises Think vs. What Actually Works With Jeff Dalton
Friday, 10 April 2026 · 2 min read · Listen to the episode ↗
In the podcast, Jeff Dalton discusses the "Agentic Gap," highlighting disparities between enterprise expectations and practical outcomes in AI applications. He emphasizes the importance of personalized coaching through Nadia, a transformative chat assistant, and advocates for precision in defining AI "agents" and evaluation metrics. The conversation also addresses the evolution of AI from deep learning to generative models, along with the challenges in measuring performance and integrating blockchain for secure data handling in financial services.
Jeff Dalton, head of AI at Valence and a professor at the University of Edinburgh, has a rich background in conversational search and is currently developing Nadia, an enterprise coaching chat assistant. His career spans nearly 20 years, beginning with a startup search engine and transitioning through research and industry, focusing on advancements in AI. At Google, he worked on health search and recognized the limitations of traditional search interfaces, leading to the development of conversational systems like Google Assistant. His research includes establishing a lab for conversational search assistants and creating evaluation benchmarks with Microsoft.
Dalton's current work emphasizes personalization, tool use, and deep research agents, highlighting the importance of precision in defining terms like "agent" and understanding historical knowledge in AI development. He advocates for human involvement in AI systems for data evaluation and encourages a "test first" approach to understand agent behavior. The conversation also addresses the evolution from deep learning to generative AI and large language models, with advancements in instruction tuning and tool integration.
Nadia aims to be more than a simple assistant, focusing on transformational coaching and personalization to improve decision-making. It aligns with organizational values and adapts over time, considering individual personalities and team dynamics. The system features a robust memory system for tracking conversations, user control over memory management, and transparency regarding what the system knows. An intelligence layer analyzes past interactions to enhance future coaching while ensuring safety within organizational constraints.
Measurement is a critical theme, with Dalton emphasizing the need for defining objective functions to evaluate performance effectively. He discusses the challenges of measuring subjective tasks and proposes grading rubrics for evaluating coaching conversations. Key evaluation pillars include coaching quality, conversation flow, and user satisfaction, with a focus on calibrating evaluators to ensure accurate assessments.
The conversation also touches on the challenges of managing multiple applications in financial services, balancing off-the-shelf solutions with bespoke systems. Dalton stresses the importance of clear requirements and specifications, advocating for sample conversations and use cases before creating prompts. He recommends starting with off-the-shelf solutions to test effectiveness before considering customization.
The future of agentic assistance is explored, focusing on agents accessing privileged knowledge about users' lives. Clear guidelines on knowledge sharing and memory retention are necessary, alongside ongoing research into knowledge federation and privacy. Nadia is presented as a hybrid model combining large language capabilities with structured task management, with concerns about granting too much freedom to task-pursuing agents.
The conversation emphasizes the necessity of evaluating outputs in systems, highlighting that not all outputs are essential. Selectivity in triggering memory is important, and structured data is needed to ensure traceability and safety. The discussion concludes with a recognition of the importance of measurement and evaluation tools tailored to specific risks and objectives, with metrics for coaching including process quality, outcome quality, and session success. Longitudinal measurement of user journeys is vital for tracking user progress and trust over time.
This summary was generated from the episode transcript and can contain mistakes.