Why Social Engineering Now Works on Machines
Tuesday, 2 December 2025 · 3 min read · Listen to the episode ↗
Ian Webster discusses the growing security challenges posed by AI agents, predicting 2026 as the "year of the agent". He emphasizes the inadequacy of traditional security methods against social engineering tactics, advocating for tools like Promptfoo to aid in securing AI interactions. The conversation highlights the increasing risks of data leakage and vulnerabilities in AI systems, urging developers to adopt dynamic security strategies rather than fixed models. Additionally, it addresses the creative nature of social engineering and its implications for AI security.
Ian Webster discusses the rapid development of AI and its implications for security, predicting that 2026 will be the "year of the agent." He highlights that security teams are struggling to keep pace with AI initiatives, often prioritizing speed over security, leading to vulnerabilities in AI agents. Traditional security measures are inadequate against social engineering and persuasion tactics, necessitating new approaches to secure AI agents.
Webster shares insights from his experience at Discord, where significant resources were dedicated to security while developing an AI agent for a large user base. He introduces Promptfoo, a tool designed to simulate adversarial conversations to test AI agents for data leaks and security vulnerabilities. Joel de la Garza from A16Z engages with Webster about the landscape of AI agents, noting various applications, including customer service bots and gaming NPCs. Webster defines an agent as a Large Language Model (LLM) that interacts with external systems via APIs, with many Fortune 10 and Fortune 50 companies planning to integrate these agents into their internal systems.
The discussion highlights that security is often an afterthought in the development cycle, with teams rushing to test before production. The ideal scenario involves equipping developers with tools like Promptfoo to actively test and receive feedback, ensuring secure systems before deployment. Recent efforts focus on integrating Promptfoo into CI/CD processes to provide security feedback directly in developers' IDEs.
Webster outlines common vulnerabilities in AI agents, including prompt injections and data leakage risks, particularly when applications connect to knowledge bases. He introduces the "lethal trifecta," identifying three critical factors that make an agent insecure: untrusted user input, access to sensitive information, and outbound communication channels. Developers are encouraged to evaluate their agents against these factors and consider limiting exposure.
The conversation addresses the subtle risks associated with data handling, where untrusted data can originate from various sources, and communication channels may inadvertently exfiltrate information. Agent security differs from traditional security due to complex interactions with diverse tools and data sources, necessitating a shift from fixed models to more dynamic approaches.
Webster points out the need for better education on acceptable security practices and the development of systems to prevent security issues. He mentions a security incident involving a SaaS provider's Gen.AI interface that allowed unauthorized data access, underscoring the importance of preventing prompt injections and access control vulnerabilities.
Human red teamers typically test security through role play, while Promptfoo simulates this process at scale, conducting numerous conversations to identify potential security flaws. The evolution of security testing is noted, with modern tools like Promptfoo generating attacks based on natural language and specific business contexts, moving away from reliance on predefined signatures.
The conversation highlights the evolving landscape of social engineering, particularly its application to AI systems. While many AI interactions are conversational, not all targets are, as some are merely API endpoints. These conversations can lead to vulnerabilities, especially concerning data leakage and access control. Human penetration testers face significant challenges, often spending extensive hours navigating these interactions, prompting a shift towards automation in security testing.
The discussion also delves into jailbreaks, described as creative expressions that exploit AI vulnerabilities. An example from a VMware researcher illustrates how informal language and emojis can bypass AI defenses. This creativity in coding parallels storytelling, emphasizing that such ingenuity is not easily accessible to the general public. The interplay of creativity and social engineering is crucial, as successful interactions with AI often hinge on exploiting access control weaknesses. The concept of "emotional fuzzing" is introduced, showcasing the surprising effectiveness of persuasion techniques on machines.
The guest shares insights from their experience at Discord, where they developed AI features for a large user base, encountering significant challenges related to security, trust, and compliance, particularly concerning risks like data exfiltration. This experience informed the development of Promptfoo, an open-source tool aimed at enhancing security evaluations. The guest notes a trend in the security industry where innovators often come from non-traditional backgrounds, influencing the creation of new security solutions.
This summary was generated from the episode transcript and can contain mistakes.