FPGA‐Based GRU Accelerator for Streaming Data Processing
Doaa K. Hameed, Abdullah M. ZyarahABSTRACT
In the evolving landscape of edge devices, unlocking artificial intelligence (AI) services via the cloud has consistently raised concerns related to energy consumption, latency, and privacy. Shifting these services locally onto edge devices powered by CPUs and GPUs have always been hindered by their limited storage and computational power. This intensifies the demand for custom‐designed AI accelerators. In this work, we propose an FPGA‐based AI accelerator capable of efficiently processing time‐series information. The proposed accelerator is built around a core algorithm modeled by a new variant of the standard gated‐recurrent unit (GRU), namely random gated‐recurrent unit (RGRU), thereby offering low latency, minimized resource utilization, and high performance compared to existing GRU accelerators. These improvements are achieved through parameter quantization, reduced computational overhead, and data localization, along with a 49 reduction in memory footprint compared to the standard GRU. The proposed design is evaluated on forecasting tasks using several standard benchmarks, including Yahoo Finance, Electricity Load, and Mackey–Glass, and demonstrates performance comparable to the standard GRU while requiring significantly fewer hardware resources.