Securing AI Agents: From Risk Assessment and Runtime Guardrails to Self-Improvement and Certification (remote)
Abstract
Foundation model powered AI agents are rapidly evolving into autonomous systems capable of reasoning, planning, and acting in complex environments. While these capabilities unlock tremendous opportunities, they also introduce fundamentally new security and safety challenges, including indirect prompt injection, agent poisoning, malicious tools, policy violations, and unsafe actions.
In this talk, I will present a unified framework for building trustworthy AI agents through four complementary pillars: risk assessment, runtime guardrails, self-improvement, and safety certification. I will first discuss emerging attack surfaces and introduce a unified platform DTap with automated red teaming techniques for systematically evaluating agent robustness across diverse adversarial scenarios with diverse environments access. I will then present a runtime agentic guardrail framework that enforces policy-compliant behavior by reasoning over an agent's execution trajectory and dynamically synthesizing verifiable safety constraints. Building on continuous evaluation, I will further discuss how red teaming and runtime monitoring can drive self-improving agents that automatically strengthen their defenses against evolving threats.
Finally, I will introduce a comprehensive benchmark for evaluating agent security across realistic web environments and safety-critical tasks, providing a foundation for standardized safety certification. Together, these advances outline a roadmap toward autonomous AI agents that are not only more capable, but also more secure, reliable, and trustworthy.
Speaker