DOI: 10.3390/computers15100667 ISSN: 2073-431X

QUINN: Quantized Inference and Dataflow Characterization for Neural Networks on RISC-V Near-Memory Architectures

Vincenzo Petrolo, Flavia Guella, Guido Masera, Maurizio Martina

Machine learning inference is moving to the edge due to latency, energy, and privacy constraints. The growing complexity of edge workloads places increasing pressure on memory hierarchies, making data movement a major limitation for both performance and energy efficiency. Fixed-function accelerators inherit the memory wall of the von Neumann model and lack the flexibility to follow evolving models. Programmable Vector Processing Unit (VPU)-based near-memory architectures alleviate this bottleneck using standard Complementary Metal-Oxide Semiconductor (CMOS) while remaining programmable, but still lack mature kernel libraries and operator-level abstractions. We address this gap with QUINN, a kernel library for integer-precision inference following quantized Open Neural Network Exchange (ONNX) operator semantics. Operator-to-dataflow mapping is well studied for conventional accelerators; we quantify it for near-memory VPUs whose vector register file is managed by explicit DMA (Direct Memory Access) transfers, using a roofline methodology combining throughput with measured off-chip traffic across dataflows, replication, and bandwidth on a Field-Programmable Gate Array (FPGA)-emulated target. Matrix multiplication enters the compute-bound regime, attaining 61% of the mix-weighted roof at the largest problem size, and one-dimensional convolution scales by 3.05× under four VPUs; the kernels sustain 5.0–51.3× the throughput of a scalar host on the same target. Operational intensity does not follow the same ordering: the scalar schedule retains higher intensity on every operator, so the throughput is bought with additional off-chip traffic in every case, though the gain exceeds the cost by at least 3.8×. QUINN thus provides a characterized operator backend, bit-exact with the ONNX semantics consumed by deployment tools.