DOI: 10.1145/3849876 ISSN: 1544-3566

SIMTSim: Exploiting Hardware Platform Parallelism for GPGPU Instruction Set Simulation

Chih-Mao Lin, Liang-Chou Chen, Yu-Yu Hsiao, Meng-Tsung Tsai, Chung-Ho Chen

The rapid proliferation of specialized GPGPU SIMT architectures for AI and data parallel workloads demands fast, functionally accurate simulation to drive ISA exploration and software stack development. However, traditional instruction set simulators face tremendous bottlenecks when modeling the massive thread concurrency of modern GPGPU workloads, severely hindering development cycles. In this paper, we introduce SIMTSim, a high performance, ahead-of-time (AOT) static binary translation framework for fast GPGPU ISA simulation. SIMTSim accelerates simulation by directly exploiting the independence of thread blocks and the lockstep execution of warps inherent to the SIMT execution model, fully leveraging the parallel execution capabilities provided by the host platform.

SIMTSim features a dual-backend architecture targeting multicore CPUs and GPUs. The CPU backend maps guest thread blocks to host worker threads and uses host SIMD instructions to evaluate warps, scaling nearly linearly with host core counts. The GPU backend employs a SIMT-on-SIMT execution strategy, mapping guest SIMT structures directly onto host GPU hardware. This approach achieves high simulation throughput, delivering near-native performance for regular kernels and generating Gemma3 1B and 4B tokens at about 2 × the native kernel time. For

gemm
, AOT translated GPU execution reaches 1, 311, 660 MIPS, compared with 18, 618 MIPS for the optimized CPU backend and 9 MIPS for the interpreter baseline. Our evaluation demonstrates that SIMTSim provides massive speedups over traditional interpretation, offering a foundational, versatile tool for rapid GPGPU ISA design, profiling, and software stack development.