AI Observability and Security for Agentic Workflows with Karthik Bharathy
Thursday, 20 March 2025 · 3 min read · Listen to the episode ↗
The discussion focuses on AI observability and security for agentic workflows, emphasizing the integration of robust governance, access controls, and monitoring to ensure data integrity and model performance. Karthik Bharathy highlights the evolution of generative AI, advocating for comprehensive observability tools and federated governance frameworks to balance standardization with flexibility. The conversation also addresses the challenges of security and data privacy in AI and blockchain applications, underlining the necessity for effective risk management and human oversight in automated systems.
Krishna Gade, CEO of Fiddler AI, engages in a discussion with Karthik Bharathy, General Manager for AI Ops and Governance at AWS SageMaker AI, on AI observability and security for agentic workflows. Karthik, with over 20 years in AI and ML, highlights the necessity of integrating security and governance into machine learning workflows, emphasizing robust data governance, access controls, and audit trails from the beginning.
The conversation underscores the significance of end-to-end observability, which involves monitoring data ingestion, lineage, and model deployment to identify drift. Despite increased automation, human oversight is crucial at key decision-making junctures, particularly in regulated industries where transparency and accountability are paramount. Karthik discusses the evolution of generative AI, noting its transition from curiosity to practical applications and large-scale deployment, while stressing the importance of security and reliability in complex workflows.
Agentic AI systems have the potential to enhance workplace productivity by automating repetitive tasks, as demonstrated by NFL Media's reduction of onboarding training time by 67% and Amazon's significant savings in software development. Enterprises are increasingly adopting agentic AI to streamline operations, as seen in Cognizant's automation of mortgage compliance workflows and Moody's multi-agent system for credit risk reports. However, challenges remain, particularly regarding security and visibility, with issues like data authenticity and unauthorized information extraction necessitating stringent access controls and data privacy measures.
Karthik emphasizes the need for comprehensive observability to monitor model performance, mentioning tools like Bedrock guardrails and Fiddler. Best practices for AI security include protecting model weights, encrypting storage, and continuous monitoring for data drifts. Application-level security practices, such as enforcing access controls and logging interaction patterns, are also vital.
The discussion advocates for a federated governance model for AI systems, balancing standardization with flexibility. Effective AI governance requires collaboration between business stakeholders and risk officers who understand the domain. Organizations must enforce standards in partnership with technical teams knowledgeable about model operations, contextualizing model toxicity scores within specific use cases.
Recommended metrics for model evaluation include documenting the model's purpose, training data, validation rules, and quality in a model card, along with explainability metrics like SHAP and LIME scores. For Generative AI, tracking additional metrics such as toxicity and fairness is essential, along with periodic evaluations against standardized datasets. Organizations face the challenge of balancing rapid AI adoption with the need for proper implementation to mitigate risks.
To manage risk, organizations should establish a range of risk values rather than a binary classification and implement additional approval workflows for higher-risk changes. Observability solutions are crucial for monitoring model drift, with alerts for minor drift and additional approvals for significant drift. Traditional metrics like ROC curves are less applicable to Generative AI, prompting a shift towards end-to-end system evaluations.
Human oversight is vital in evaluating automated system outputs, especially during early deployment phases, and controls on data access and output generation are necessary. Implementing a pause or kill switch for models deviating from known patterns is recommended. Variability in measurement and security controls across industries necessitates adherence to specific policies and regulations, such as the EU Act and ISO 42001.
Robust authentication and authorization controls are essential, emphasizing re-authentication to ensure only authorized personnel access specific jobs. The principle of least privilege should be applied. The conversation also touches on best practices for code migration and the importance of utilizing tools like Amazon Q for workflow assessment and security considerations.
Karthik stresses the importance of starting with high-quality data and establishing a robust data infrastructure to support effective AI models. He advocates for focusing on high-value use cases to prototype solutions that address specific business problems. Governance frameworks are crucial as organizations scale, necessitating the establishment of approval workflows and team training to build internal expertise.
Karthik warns against the pitfalls of rushing implementations without ensuring data quality and stakeholder engagement, emphasizing the need for a deep understanding of AI models to avoid maintenance and deployment issues. He notes the ongoing demand for improved performance, robustness, security, and cost-effectiveness in AI agents, as well as the evolving landscape of AI applications supported by partnerships and diverse model offerings.
This summary was generated from the episode transcript and can contain mistakes.