Beyond Vibes: Evaluation Strategies for Safe Multi-Turn AI Agents

Tuesday, August 11, 2026
12:00 PM - 12:30 PM
AI Risk Summit Tech Track (Salon II)

About This Session

Evaluating whether an AI agent completes a task is hard. Evaluating whether it does so safely is harder — and most teams have no systematic way to catch failures before users do. What breaks when you try to measure safe behavior across multi-step, non-deterministic agent workflows? Why aren’t classic eval pipelines built for this? And what practical evaluation strategies actually work? This talk tackles all three, offering concrete patterns — from trajectory-level assertions to adversarial scenario generation to safety-aware scoring rubrics — for teams shipping agents today. Drawing on experience building LLM evaluation frameworks and production safety systems at scale, we go beyond surface-level pass/fail metrics to show how you can build real, repeatable confidence that your agent is behaving as intended.

Speaker

Eti Rastogi

Eti Rastogi

Sr. Applied Scientist - Amazon

Eti Rastogi is an Applied Scientist at Amazon AWS working on AI safety, LLMs, and multimodal systems. She leads the science direction for Amazon Bedrock Guardrails — the AI safety system protecting tens of thousands of enterprises globally. Previously, she was the first AI hire at DeepScribe, where she built the entire AI function from scratch to power documentation for over 3 million cancer care visits annually. At Scale AI, she built evaluation systems for frontier models used by hundreds of millions of people.