Mohit Gupta
Session Speaker
I am a Staff GPU Performance Verification Engineer at Intel, specializing in GPU microarchitectural performance analysis, verification, benchmarking, and optimization. I have over five years of pre-silicon design verification experience and more than three years of experience focused on GPU performance verification. At Intel, I lead performance verification and qualification of GPU compute subsystems, with expertise in front-end dispatch pipelines, Load-Store Cache (LSC) units, HPC workloads, performance profiling, and bottleneck analysis. Previously, I have worked as a Senior Design Verification Engineer at Qualcomm Technologies, where I have developed scalable UVM-based verification environments, GPU verification infrastructure, and performance validation methodologies. My technical expertise includes GPU architecture, performance analysis and tuning, profiling and benchmarking, SystemVerilog, Verilog, Python, UVM, and constrained-random verification.I holds a Master of Science in Electrical Engineering from San Jose State University and a Bachelor of Technology in Electronics & Communication Engineering from Kurukshetra University
Title: Stress Testing Compute-Bound vs Memory-Bound Workloads in GPU Microarchitecture Validation - Abstract. Modern GPU architectures must accommodate increasingly diversecomputational workloads ranging from artificial intelligence inference to high-performancescientific computing, each presenting distinct microarchitectural stress patterns that challengetraditional validation methodologies. This paper presents a comprehensive framework forsystematic stress testing of GPU microarchitectures through targeted differentiation betweencompute-bound and memory-bound workloads, addressing significant gaps in conventionalverification approaches that rely on uniform testing strategies inadequate for contemporaryheterogeneous computing demands. The framework establishes quantitative classificationcriteria for workload categorization based on resource utilization patterns, implementsrepresentative benchmark suites spanning realistic application domains, and providesperformance analysis techniques enabling early identification of architectural bottlenecksduring pre-silicon validation phases. Through systematic stress testing across diverse scenarios,including dense linear algebra operations, irregular memory access patterns, and mixedworkload interactions, the methodology exposes previously undetected corner cases andsubsystem interaction effects that traditional approaches consistently miss. Results demonstratesubstantial validation improvements: 60% better bottleneck detection compared to traditionalapproaches, 40% reduction in post-silicon debugging time, and 25% improvement inarchitectural optimization accuracy. Industry applications confirm reduced development cyclesand enhanced product performance validation through comprehensive stress testingmethodologies that guide informed design decisions based on realistic workloadcharacterization rather than synthetic benchmarks.