DOI: 10.3390/computers15100673 ISSN: 2073-431X

GPU-Parallelization of the Numerical Solution of Large-Scale Transient Partial Differential Equations

Dániel Koics, Endre Kovács, Olivér Hornyák

Efficient and scalable numerical methods for time-dependent partial differential equations are of central importance. This study extends our previous examinations of OpenCL-based parallel implementation of three explicit finite difference schemes: the one-stage Constant Neighbor (CNe) method, its two-stage predictor–corrector variant (CpC), and the well-known Euler’s method. The CNe and CpC schemes are unconditionally stable for the diffusion equation, and the Euler scheme serves as a reference method. However, the investigations are now extended to significantly larger spatial systems—up to 1.6 billion nodes—and refined timescales. One- and two-dimensional initial value problems are solved with recently published non-trivial analytical reference solutions. Beyond general benchmarking of error versus execution time and execution time versus grid size, two focused analyses are performed: (i) assessment of GPU performance depletion as the number of spatial grid points approaches extreme scales, and (ii) detailed investigation of the error convergence behavior of the CpC method under intensive timestep refinement. Results confirm the near grid-size-independent runtime characteristics of the parallel implementations within practical limits, while the onset of GPU resource saturation is also identified at very large problem sizes. For a 500-timestep calculation, we can conclude that GPU parallelization starts to be competitive above 2–3 thousand nodes, and the gain in computational time reaches a factor of 5–40 at around 109 nodes. These findings highlight not only the numerical characteristics of the CpC scheme but also important practical memory constraints that arise in large-scale, GPU-accelerated explicit diffusion simulations.