Rethinking How we Evaluate Security Agents for Real-World Use
About This Session
Security agents are gaining momentum across industry, but the way we evaluate them remains rooted in narrow, outcome-only benchmarks. These evaluations tell us whether an agent produced a correct answer, but not “how” it arrived there or whether that behavior will remain stable once deployed.
In practice, enterprise security is not a sequence of isolated tasks. It is a connected, end-to-end workflow that follows a find → confirm exploit → patch → validate loop. Agents that perform well on task-specific benchmarks often fail in these multi-stage settings due to contextual loss and brittle transitions across steps.
This talk introduces a practical framework for evaluating security agents by mapping agentic capabilities (planning, reasoning, memory, perception, tool use) to security functions (reconnaissance, exploit confirmation, root-cause analysis, patching, validation) across the full lifecycle.
We also share insights from our large-scale survey of existing agentic systems and presents a lightweight, unified end-to-end scoring perspective that teams can use to assess an agent’s readiness for real operational environments.
In practice, enterprise security is not a sequence of isolated tasks. It is a connected, end-to-end workflow that follows a find → confirm exploit → patch → validate loop. Agents that perform well on task-specific benchmarks often fail in these multi-stage settings due to contextual loss and brittle transitions across steps.
This talk introduces a practical framework for evaluating security agents by mapping agentic capabilities (planning, reasoning, memory, perception, tool use) to security functions (reconnaissance, exploit confirmation, root-cause analysis, patching, validation) across the full lifecycle.
We also share insights from our large-scale survey of existing agentic systems and presents a lightweight, unified end-to-end scoring perspective that teams can use to assess an agent’s readiness for real operational environments.
Speaker
Mudita Khurana
Staff Security Engineer - Airbnb
Mudita Khurana is a Tech Lead at Airbnb, where she builds scalable security tooling and automation across the software development lifecycle. Previously at Meta, she drove key initiatives in product security, including bug bounty strategy, privacy-focused reviews, and automated vulnerability detection through static and hybrid analysis. Her work focuses on advancing security automation through robust workflows, agentic systems, and scalable system design. Mudita also serves on program committees for several top-tier security conferences and regularly contributes to the community through research, talks, and industry collaborations.