Breaking the Sound Barrier: End-to-End Red Teaming for AI Voice Agents
About This Session
AI voice agents are a lot like their text-based counterparts. Helpful, deliberate, and vulnerable. As voice becomes a more prominent modality for user interaction and enterprise product capabilities, safety and security testing lags far behind. Voice introduces novel attack surfaces and the practical challenges of automating audio make it difficult to efficiently scale testing.
Voice agents built on speech-to-text pipelines inherit the vulnerabilities of their text-based counterparts, but introduce new attack surfaces at the transcription layer: acoustic adversarial inputs, prosody manipulation, and exploitable gaps between what a human hears and what the system transcribes. Native audio models present a different risk profile, processing speech without an intermediate text representation and creating attack vectors at the layer of sound.
Automating voice agent testing is practically harder than text. Full pipeline testing requires generating adversarial audio inputs programmatically, handling spoken latency and interruptions, and achieving meaningful coverage across languages and dialects. These challenges contribute to an immature safety and security landscape across a quickly emerging modality.
Attendees will leave with an understanding of the voice agent risk landscape, how attackers approach reconnaissance and exploitation, and practical guidance for building safety and security audio testing programs that scale.
Voice agents built on speech-to-text pipelines inherit the vulnerabilities of their text-based counterparts, but introduce new attack surfaces at the transcription layer: acoustic adversarial inputs, prosody manipulation, and exploitable gaps between what a human hears and what the system transcribes. Native audio models present a different risk profile, processing speech without an intermediate text representation and creating attack vectors at the layer of sound.
Automating voice agent testing is practically harder than text. Full pipeline testing requires generating adversarial audio inputs programmatically, handling spoken latency and interruptions, and achieving meaningful coverage across languages and dialects. These challenges contribute to an immature safety and security landscape across a quickly emerging modality.
Attendees will leave with an understanding of the voice agent risk landscape, how attackers approach reconnaissance and exploitation, and practical guidance for building safety and security audio testing programs that scale.
Speaker
Matt Fiedler
Sr. Product Manager, AI Agent Security - Check Point
Matt Fiedler is a Senior Product Manager and member of the Office of the CTO at Check Point, where he leads the AI red teaming product and professional services. He joined Check Point through the acquisition of Lakera, where he worked across the AI safety and security portfolio. His work focuses on the trust, safety, and security of agentic AI systems. He holds an M.S. in Computer Science from Johns Hopkins and played professional baseball in the St. Louis Cardinals organization.