Rethinking How we Evaluate Security Agents for Real-World Use

Wednesday, August 12, 2026
3:30 PM - 4:00 PM
AI Risk Summit Tech Track (Salon II)

About This Session

Security agents are gaining momentum across industry, but the way we evaluate them remains rooted in narrow, outcome-only benchmarks. These evaluations tell us whether an agent produced a correct answer, but not “how” it arrived there or whether that behavior will remain stable once deployed.

In practice, enterprise security is not a sequence of isolated tasks. It is a connected, end-to-end workflow that follows a find → confirm exploit → patch → validate loop. Agents that perform well on task-specific benchmarks often fail in these multi-stage settings due to contextual loss and brittle transitions across steps.

This talk introduces a practical framework for evaluating security agents by mapping agentic capabilities (planning, reasoning, memory, perception, tool use) to security functions (reconnaissance, exploit confirmation, root-cause analysis, patching, validation) across the full lifecycle.

We also share insights from our large-scale survey of existing agentic systems and presents a lightweight, unified end-to-end scoring perspective that teams can use to assess an agent’s readiness for real operational environments.

Speaker

Mudita Khurana

Mudita Khurana

Staff Security Engineer - Airbnb

Mudita Khurana is a Tech Lead at Airbnb, where she builds scalable security tooling and automation across the software development lifecycle. Previously at Meta, she drove key initiatives in product security, including bug bounty strategy, privacy-focused reviews, and automated vulnerability detection through static and hybrid analysis. Her work focuses on advancing security automation through robust workflows, agentic systems, and scalable system design. Mudita also serves on program committees for several top-tier security conferences and regularly contributes to the community through research, talks, and industry collaborations.