FPGA-Based Acceleration of a Multiscale Decision Fusion CNN for Radar-Based Object Classification
Turki Aldosari, Abdullah Alghaihab, Saleh AlshebeiliThe increasing use of unmanned aerial vehicles (UAVs) and other moving objects in civilian environments has created a growing need for reliable radar-based object classification systems, while deep learning (DL) models achieve high classification accuracy on range–Doppler (RD) radar data, their deployment on resource-constrained embedded platforms remains challenging due to computational and latency constraints. This work addresses that gap as a hardware–software co-design problem: the contribution is the system-level optimization and characterization of the deployment rather than the design of a new network architecture. This paper presents the deployment and system-level characterization of an adopted multiscale decision fusion convolutional neural network (MDF-CNN) for radar-based object classification on the Xilinx Kria KR260 platform. The network processes preprocessed stacked RD frames of size 11×61×3 and classifies targets into cars, drones, and people using the publicly available RDRD dataset. The Floating-Point (FP32) reference model achieves 97.06% classification accuracy. For efficient embedded deployment, the model is quantized to 8-bit integer (INT8) precision using the Vitis AI framework and executed on the deep learning processing unit (DPU). The INT8 FPGA (DPU) implementation achieves 96.88% accuracy with an average classifier-stage latency of 9.24 ms per preprocessed input tensor and a throughput of 108.21 inferences per second, corresponding to a 6.2× latency reduction relative to the FP32 central processing unit (CPU) baseline. The measured interval includes DPU execution and the tensor transfers between double data rate (DDR) memory and the DPU, but excludes radar acquisition, RD-map generation, target-centered cropping, temporal stacking, per-sample z-score normalization, and ARM-side INT8 input scaling. At fixed INT8 precision, the INT8 CPU implementation achieved a lower raw latency of 7.957 ms per preprocessed input tensor compared with 9.24 ms for the DPU. Power measurements show an average consumption of 4.14 W, resulting in an energy cost of 38.25 mJ per classifier inference. These figures compare two deployment configurations that differ in hardware, numerical precision, software runtime, and power-measurement boundary, and are therefore reported as a system-level comparison rather than as an attribution to any single factor. These results characterize low-latency and energy-efficient deployed MDF-CNN classifier execution on the KR260 using preprocessed radar tensors; investigation of the implementation and timing of the complete online radar-to-classification pipeline remain directions for future work.