Bird’s-Eye-View Road Occupancy Prediction for Autonomous Driving: A Survey of Representations, Methods, and Benchmarks
Abdelrahman S. Heikal, Mostafa Farouk Senussi, Ahmed Salem, Hyun-Soo KangBird’s-eye-view (BEV) perception has become the dominant paradigm for camera-centric scene understanding in autonomous driving, as well as road occupancy prediction, which involves the dense estimation of which regions of space are occupied and by what has emerged as its most expressive form. Between 2020 and 2026, the field underwent three overlapping transitions: from two-dimensional BEV semantic map segmentation to dense three-dimensional voxel-based 3D semantic occupancy, catalyzed by the 2022 industrial adoption of “occupancy networks,” and, most recently, to efficient, generative, and four-dimensional forecasting formulations. This survey organizes the literature along six orthogonal axes output representation, view-transformation mechanism, input modality, supervision paradigm, temporal scope, and efficiency strategy and uses the representation lineage as a primary spine connecting the 2020 BEV-segmentation works to the 2026 Gaussian and 4D frontier. Alongside the ego-centric mainstream, we review the parallel multi-view and infrastructure-side lineage from multi-view pedestrian occupancy to roadside traffic occupancy, which shares the BEV occupancy-map output and contributes generalization tools the ego-centric thread has yet to absorb. We review the canonical methods at each stage, summarize the standard datasets (CARLA, GMVD, MultiviewX, WildTrack, nuScenes, SemanticKITTI, Occ3D, OpenOccupancy) and evaluation metrics (MODA, mIoU, RayIoU, RayPQ), and consolidate reported results on the Occ3D-nuScenes benchmark into a single comparison. We close by identifying open problems in label efficiency, robustness, temporal forecasting, and deployment. Our intent is to bridge the historically separate BEV-segmentation and 3D-occupancy literatures within a single taxonomy.