Anudeep Rentala

Title of the Talk: Beyond the Happy Path: Operating Distributed Systems at Millions of Queries per Second

Abstract :
At millions of queries per second, distributed systems fail in ways that rarely appear in architecture diagrams. Small latency regressions become capacity crises, retries amplify overload, downstream failures cascade across services, and even routine deployments can threaten availability. This talk presents practical principles for building and operating high-throughput systems where reliability must be engineered continuously rather than added after launch. Drawing on experience operating critical services at more than three million peak QPS, Anudeep Rentala examines how to define meaningful service-level indicators, control retry amplification, isolate failures, manage hot spots, degrade gracefully, and execute safe production rollouts. The talk also explores the organizational side of reliability: turning incidents into durable engineering improvements and coordinating changes across multiple teams without slowing product development. It offers a practical framework for designing systems that remain observable, scalable, and dependable as traffic and organisational complexity grow.

Bio :
Anudeep Rentala is a Senior Software Engineer specializing in large-scale distributed systems, machine-learning serving infrastructure, and production reliability. He builds ranking, experimentation, and critical backend infrastructure for Roblox’s Avatar Marketplace, serving more than 100 million monthly users. He has led reliability improvements across services handling millions of requests per second, increasing availability from approximately 99% to 99.995%, reducing incidents by 90%, and delivering millions of dollars in annual infrastructure savings. Previously at Splunk, he built observability and infrastructure-monitoring products that helped organizations operate cloud systems reliably. He is also a named inventor on a granted U.S. machine-learning patent and has delivered company-wide technical talks on distributed systems, applied AI, and production reliability.