A Real-Time Single Image Haze Removal Accelerator for Increasingly Automated Vehicles
Yifu Zhu, Yanjie Tan, Feiteng Nie, Zhaoyang Huang, Huailiang TanAutonomous vehicles at L2 and above are increasingly relying on stereo vision systems, where haze removal is critical to detect obstacles hidden in fog. Existing image haze removal techniques have low processing speed and high resource consumption, restricting their application scope in practice. In this work, we propose a hardware-software co-design solution for haze removal. It fully decouples the calculation of the two main parameters, i.e., atmospheric light and transmission, in the dehazing process. By eliminating the data dependency, parallelism in hardware acceleration is enhanced. Furthermore, in replacement of the conventional global homogeneous atmospheric light computation, we report a chunk-based heterogeneous method to reduce cache overhead. Our approach is implemented on both GPU and FPGA, compared against six state-of-the-art (SOTA) works for image haze removal. Evaluation using test sets of real-world foggy driving scenarios shows that our object detection accuracy is over 88%, 9.5%-47.4% better than the SOTA works with neural networks (NN) on GPU, 42.2% better than the non-NN SOTA on GPU, and 25.9%-52.2% better than the SOTA works on FPGA. The processing speed varies with image resolution and our improvement is generally even more at higher resolution. In FPGA implementation, our approach is 29.7% faster than the fastest SOTA at the lowest resolution of 360p. We have the lowest overall resource consumption, where the bottleneck BRAM usage is reduced by over 70%. In GPU implementation, our approach is 2-3 orders of magnitude faster than the NN-based SOTA works, and saves two orders of magnitude in memory, from about 10GB to hundreds of MB. The non-NN SOTA on GPU also consumes hundreds of MB in memory, and we are 31.4% to 66.5% faster than it, at different resolutions. Our power consumption is 5.7W and 53W, the lowest in the FPGA and GPU category, respectively. To compare between ourselves in timing, both take several milliseconds. The GPU solution is slightly faster and more scalable, yet with fluctuation of -5.3% to +22.2%. The FPGA solution has circuit-level timing determinism at nanosecond, hence suitable for hard real-time applications.