PodBrowser
AI Explained

Governing AI That Keeps Evolving With Maryam Ashoori (VP of Product and Engineering at IBM watsonx.governance)

Thursday, 6 August 2026 · 4 min read · Listen to the episode ↗

Maryam Ashoori, VP of Product and Engineering at IBM watsonx.governance, draws on two decades of experience with multi-agent systems to argue that the core governance, risk, and compliance challenges have not fundamentally changed even as agentic AI has surged into enterprise production.

Maryam Ashoori, VP of Product and Engineering at IBM watsonx.governance, worked on multi-agent systems 20 years ago for her master's dissertation and argues the core governance, risk, and compliance challenges have not fundamentally changed since then, even as specific risks have evolved. She traces the post-ChatGPT enterprise cycle through three phases: an initial wave of narrow use cases around summarization, extraction, and code generation; a second phase where broad generative AI deployment failed to demonstrate ROI; and a mid-2024 shift to agents, driven by board mandates to have agents in production by end of 2024 without clarity on what problem they would solve.

The primary blocker to production deployment of agentic AI is not proof of value from experimentation but uncertainty about whether proper instrumentation, guardrails, and processes exist to manage agent autonomy. Ashoori cites a figure that 75 percent of business executives are not confident they would pass an independent AI audit within 90 days, and many of those executives already have agents running in production. She also cites an estimate that a 20 billion dollar enterprise with weak governance loses approximately 70 million dollars a year from preventable oversight failures.

Ashoori frames trust in AI systems as resting on three layers: visibility, control, and accountability. She is explicit that control without accountability is only good intention and requires enforcement mechanisms. Accountability is the number one challenge cited in her customer advisory boards because agentic decisions involve multiple simultaneous parties including business units, model providers, third-party agent providers, tool providers, and end users, and no single company can resolve accountability to external parties without regulation catching up. On agent identity, she says the market collectively does not yet understand who the accountable party is when an agent fails, and that in the near term the person applying the agent is responsible and the agent inherits that person's access controls.

The AI control plane emerged as a recognized market category within approximately the last six months as of the episode. Ashoori defines a control plane as needing to define controls, implement controls, enforce controls, and close the loop by tracking impact and adjusting. Controls derive from four sources: corporate policies, enterprise AI risk, regulatory compliance, and operational considerations such as token costs. She distinguishes between security enforcement tools covering risks such as the OWASP top 10 agentic application risks and observability dashboards covering AI-layer concerns such as monitoring for sensitive information in outputs. Her view is that enterprises need multiple specialized observability platforms but a single unified governance console, and that governing AI is best handled by a third party rather than the party building it, analogous to SOC 2 certification.

Traditional governance built on documentation, approvals, and quarterly reviews was not designed for continuous delivery. Ashoori argues governance must be present at the earliest moment a developer conceives of a use case, before anything is built, because runtime tracing alone only shows what already happened and may be too late to prevent harm such as data leakage. IBM watsonx.governance performs use case similarity analysis at build time to surface existing approved applications with their constraints, and automatically routes high-risk use cases to risk and compliance teams. She describes three layers of agent evaluation: in-node evaluation during execution, evaluation over time to detect drift, and offline evaluation at build time. A core unsolved industry problem she identifies is not knowing what good looks like or what the ground truth is for agentic trajectories.

Ashoori describes governance structurally as a set of relationships connecting AI assets to use cases, use cases to risk, risk to controls, controls to metrics, and metrics to business objectives. She argues most AI governance solutions focus on AI assets in isolation from enterprise operations, which is where they fall short. IBM watsonx.governance uses a governance graph to trace those relationships all the way back to business initiatives, allowing organizations to identify exactly which KPIs and financial outcomes are affected when a control is breached.

On the competitive landscape, she argues models are becoming increasingly interchangeable while enterprises sitting on specialized proprietary data have a durable advantage because most LLMs are trained on public data. She says operationalizing an AI agent is approximately 90 percent more difficult than building one and that almost all agents in existence are not designed for production. Orchestration costs across multiple specialized agents can exceed the cost of running a single large model, making the right architecture entirely use-case dependent. The most popular enterprise AI use case remains content-grounded question and answering, and she states that system reliability is more important than model accuracy. She prefers the phrase human in the lead over human in the loop because it positions humans as active orchestrators rather than passive checkpoints.

This summary was generated from the episode transcript and can contain mistakes.