Global Conference on Intelligent Software Architecture, AI/ML Engineering & Cloud Computing

Theme: "Bridging Intelligent Software Architecture, AI/ML Engineering, and Cloud-Native Technologies for the Future"

12-13, November 2026 Seri Pacific Hotel Kuala Lumpur, Kuala Lumpur, Malaysia
Back to conference
Mohit Gupta
Featured Speaker

Mohit Gupta

Session Speaker

USA

Biography

I am a Staff GPU Performance Verification Engineer at Intel, specializing in GPU microarchitectural performance analysis, verification, benchmarking, and optimization. I have over five years of pre-silicon design verification experience and more than three years of experience focused on GPU performance verification. At Intel, I lead performance verification and qualification of GPU compute subsystems, with expertise in front-end dispatch pipelines, Load-Store Cache (LSC) units, HPC workloads, performance profiling, and bottleneck analysis. Previously, I have worked as a Senior Design Verification Engineer at Qualcomm Technologies, where I have developed scalable UVM-based verification environments, GPU verification infrastructure, and performance validation methodologies. My technical expertise includes GPU architecture, performance analysis and tuning, profiling and benchmarking, SystemVerilog, Verilog, Python, UVM, and constrained-random verification.I holds a Master of Science in Electrical Engineering from San Jose State University and a Bachelor of Technology in Electronics & Communication Engineering from Kurukshetra University

Abstract Title

Title: Stress Testing Compute-Bound vs Memory-Bound Workloads in GPU Microarchitecture Validation - Abstract. Modern GPU architectures must accommodate increasingly diversecomputational workloads ranging from artificial intelligence inference to high-performancescientific computing, each presenting distinct microarchitectural stress patterns that challengetraditional validation methodologies. This paper presents a comprehensive framework forsystematic stress testing of GPU microarchitectures through targeted differentiation betweencompute-bound and memory-bound workloads, addressing significant gaps in conventionalverification approaches that rely on uniform testing strategies inadequate for contemporaryheterogeneous computing demands. The framework establishes quantitative classificationcriteria for workload categorization based on resource utilization patterns, implementsrepresentative benchmark suites spanning realistic application domains, and providesperformance analysis techniques enabling early identification of architectural bottlenecksduring pre-silicon validation phases. Through systematic stress testing across diverse scenarios,including dense linear algebra operations, irregular memory access patterns, and mixedworkload interactions, the methodology exposes previously undetected corner cases andsubsystem interaction effects that traditional approaches consistently miss. Results demonstratesubstantial validation improvements: 60% better bottleneck detection compared to traditionalapproaches, 40% reduction in post-silicon debugging time, and 25% improvement inarchitectural optimization accuracy. Industry applications confirm reduced development cyclesand enhanced product performance validation through comprehensive stress testingmethodologies that guide informed design decisions based on realistic workloadcharacterization rather than synthetic benchmarks.