PodBrowser
Practical AI

Controlling AI Models from the Inside

Tuesday, 20 January 2026 · 3 min read · Listen to the episode ↗

The discussion emphasizes the importance of controlling AI models from within to enhance safety and security. Key insights include distinguishing between "AI for security" and "security for AI," the need for proactive risk mitigation through model transparency, and the development of customizable safety measures. Additionally, the conversation highlights how advancements in interpretability and monitoring can prevent harmful outputs, laying the groundwork for future innovations in AI and model-native safety.

Daniel Wightnack and Chris Benson engage in a discussion on AI safety and security with guest Ali Kachri, who differentiates between "AI for security," which addresses security challenges using AI, and "security for AI," which focuses on protecting AI models. As AI models become more integrated into technology, they present new security challenges, including the risk of generating harmful content.

Ali emphasizes the importance of safety in AI, noting that models can produce inappropriate material and encourage harmful behaviors. Current safety measures often analyze inputs and outputs reactively, which can be insufficient as harmful content may be generated before detection. He highlights the challenges of model transparency, stating that without understanding AI models' internal workings, preventing malicious outputs becomes difficult. Users should identify undesirable content categories and develop context-specific risk mitigation strategies.

The conversation also explores terminology related to AI safety, such as "guardrails" and "safeguarding," which refer to protective measures for managing AI interactions. Companies have implemented guard models to filter prompts and responses, ensuring compliance with safety standards. The field of interpretability in AI is growing, focusing on understanding model outputs and modifying behavior at the source.

Ali discusses the complexities AI tools introduce in collaboration, emphasizing that clarity is more important than speed in bridging the gap between idea generation and execution. He mentions innovations like Miro's workspace, which enhances teamwork through AI-assisted features. Interpretability is crucial for understanding model decisions, especially in sensitive applications like insurance and healthcare.

The discussion addresses the relationship between safety and AI model outputs, particularly concerning jailbreaks. The speakers advocate for a proactive understanding of model behavior rather than treating it as a black box, emphasizing the need for intervention methods to control and prevent undesirable outputs. They propose monitoring activated subspaces during output generation to differentiate relevant from irrelevant areas based on application context.

Using an apartment building analogy, they stress the need for constant visibility into model operations to identify issues preemptively. They propose a new layer of safety that gathers intelligence for decision-making and risk mitigation without requiring customers to train new models. Instead, a safety module can be added to existing models for low-friction integration.

Customization is highlighted as essential, as different companies may require tailored safety measures. The distinction between external guardrails and internal model modifications is clarified, with current filters often proving ineffective due to high computational costs. A research breakthrough enabling the development of safety measures at a lower cost is mentioned.

The speakers acknowledge limitations of current safety models on edge devices, which often lack necessary memory and guardrails. They argue that while external guardrails may provide some safety, they cannot guarantee complete security within the model. Instrumentation is crucial for identifying toxicity without needing to predict every possible input, and leveraging past examples can help gauge potential problems.

Ali draws an analogy to EEGs in neuroscience, suggesting a need for hybrid approaches to enhance security models. A defense-in-depth strategy is essential, as no single product can address all security issues. In practical applications, such as customer service bots, rules can be composed to block users based on specific criteria, improving AI systems' safety profiles.

Ali discusses the future of model security, distinguishing between build time safety and runtime safety. He aspires to develop model-native safety, which is vital for broader adoption in settings like healthcare. Public LLMs face challenges due to data concerns, particularly with PII data, limiting ecosystem access. He envisions creating a de facto model safety layer applicable to any model, establishing a standard for model safety.

This summary was generated from the episode transcript and can contain mistakes.