Santosh Appachu Devanira Poovaiah
Session Speaker
I am an experienced engineer with 7 years in SoC design and verification, currently focused on full-chip coherency verification of advanced, high-performance computing systems. At NVIDIA, I actively contribute to validating flagship architectures such as Grace Hopper and Blackwell GPUs, which power DGX-class platforms used across AI, data center, automotive, and HPC applications. I have a strong track record of verifying silicon tailored for AI acceleration, working closely with industry-leading designs that push the boundaries of compute, memory, and interconnect complexity. I hold a Master’s degree in Computer Engineering from the University of Southern California. Prior to NVIDIA, I developed security-focused firmware for Robert Bosch, gaining hands-on experience bridging low-level software and hardware security. Well-versed in modern computer architecture, I possess deep expertise in cache coherence protocols, multi-agent interconnect systems, and memory consistency models. My work centers on verifying complex heterogeneous SoC architectures, including CPU-GPU interactions under parallel, high-throughput workloads. I am passionate about enabling robust, scalable, and high-efficiency hardware that powers next-generation AI training and inference systems, continuously driven by the challenge of verifying correctness and security in large, tightly integrated computing platforms.
The AI Performance Illusion: A Verification Perspective on GPU, Cloud, and Deployment Variability Enterprises often assume that AI performance is primarily determined by the model and the GPU. In reality, production deployments repeatedly reveal an “AI performance illusion”: identical models can behave very differently across GPU types, cloud platforms, and deployment environments, even when the hardware appears similar. These performance gaps are frequently misunderstood as model issues, while the root causes originate from lower-level execution and infrastructure layers such as memory movement, kernel scheduling, runtime configuration, data pipelines, container environments, and orchestration behavior. This talk introduces a verification-driven perspective for diagnosing and preventing performance variability in modern AI systems. Instead of relying on trial-and-error tuning, we treat AI performance as a system property that must be validated using repeatable measurement methodology, controlled experiments, and infrastructure-aware observability. The session breaks down real-world causes of GPU underutilization, throughput collapse, and latency spikes, covering CPU - GPU interaction bottlenecks, dynamic batching behavior, precision modes, multi-process interference, software stack inconsistencies, and cloud deployment drift. Attendees will leave with a practical blueprint for verifying AI performance across environments, ensuring reproducibility, and building deployment pipelines that deliver predictable throughput, stable latency curves, and cost-efficient scaling.