Agentic Infrastructure Risk: Securing Autonomous AI Systems in Production
About This Session
AI systems are rapidly evolving from passive models into autonomous agents capable of executing complex workflows across APIs, services, and cloud infrastructure. While this unlocks unprecedented automation, it also introduces a new and largely unaddressed category of risk.
In this talk, we explore agentic infrastructure risk - the failure modes, security vulnerabilities, and operational hazards that emerge when AI systems are allowed to take real-world actions. Drawing from experience building large-scale infrastructure at Meta, we examine how agent runtimes interact with APIs, tools, and micro-services, and why traditional security and reliability models fail in non-deterministic systems.
We will analyze critical emerging risks, including hallucinated actions, cascading retries, privilege escalation through tool misuse, and unbounded execution loops that can trigger system-wide incidents or cost explosions. Unlike traditional software failures, these risks are amplified by autonomy and feedback loops.
The session will introduce practical mitigation strategies:
* Guardrail architectures and action validation pipelines
* Policy-driven execution controls and zero-trust action models
* Observability for reasoning-driven, non-deterministic systems
* Cost containment and runtime safety mechanisms
Attendees will leave with a concrete, production-ready blueprint for designing AI systems that are secure, observable, and governed by design - enabling organizations to deploy autonomous agents without compromising safety or control.
Key Takeaways
- Identify new classes of AI risk introduced by autonomous agents
- Understand how agent failures differ from traditional system failures
- Learn to design guardrails and zero-trust execution layers
- Build observability for decision-making systems
Apply cost, safety, and governance controls in production
In this talk, we explore agentic infrastructure risk - the failure modes, security vulnerabilities, and operational hazards that emerge when AI systems are allowed to take real-world actions. Drawing from experience building large-scale infrastructure at Meta, we examine how agent runtimes interact with APIs, tools, and micro-services, and why traditional security and reliability models fail in non-deterministic systems.
We will analyze critical emerging risks, including hallucinated actions, cascading retries, privilege escalation through tool misuse, and unbounded execution loops that can trigger system-wide incidents or cost explosions. Unlike traditional software failures, these risks are amplified by autonomy and feedback loops.
The session will introduce practical mitigation strategies:
* Guardrail architectures and action validation pipelines
* Policy-driven execution controls and zero-trust action models
* Observability for reasoning-driven, non-deterministic systems
* Cost containment and runtime safety mechanisms
Attendees will leave with a concrete, production-ready blueprint for designing AI systems that are secure, observable, and governed by design - enabling organizations to deploy autonomous agents without compromising safety or control.
Key Takeaways
- Identify new classes of AI risk introduced by autonomous agents
- Understand how agent failures differ from traditional system failures
- Learn to design guardrails and zero-trust execution layers
- Build observability for decision-making systems
Apply cost, safety, and governance controls in production
Speaker
Nishant Gupta
Staff Software Engineer, Tech Lead - Meta (Meta SuperIntelligence Lab)
Nishant Gupta is a Staff Software Engineer and Researcher at Meta, specializing in large-scale distributed systems and AI infrastructure. At Meta Superintelligence Labs, his work focuses on building reliable and secure agentic systems that operate across APIs, services, and cloud environments, with an emphasis on evaluation, guardrails, observability, and alignment with real-world outcomes.
He has also led elastic compute infrastructure managing approximately 30% of Meta’s capacity, delivering significant efficiency gains. His research spans safe resource oversubscription and production AI infrastructure and has received more than 90 citations.
He has also led elastic compute infrastructure managing approximately 30% of Meta’s capacity, delivering significant efficiency gains. His research spans safe resource oversubscription and production AI infrastructure and has received more than 90 citations.