Compressing Transformer-Based Sensor Fusion Models for Autonomous Driving Using Curriculum-Based Multi-Task Knowledge Distillation
Dhieddine Barhoumi, Stefan Hensel, Marin B. MarinovMultimodal sensor fusion architectures such as TransFuser achieve strong waypoint-based driving performance by fusing RGB panoramas and LiDAR bird’s-eye-view (BEV) representations through transformer attention, but their dual backbones and quadratic-cost fusion blocks preclude deployment on resource-constrained platforms. This work proposes LightFuser, a hybrid Performer–Transformer fusion architecture that retains multimodal reasoning at substantially reduced cost. LightFuser halves backbone parameters and FLOPs, applies Performer attention with FAVOR+ at the early, high-resolution fusion stage to obtain linear time and memory complexity, and reserves exact Transformer attention for the mid and late stages, where token counts are small and semantics are richer. A multi-task knowledge distillation framework transfers the decision quality of a high-capacity TransFuser teacher to the lightweight student, organized as a task-ordered curriculum: supervision starts with waypoint control and successively adds depth, semantic segmentation, and BEV objectives. On the CARLA Leaderboard, LightFuser attains a Driving Score of 42.4 (teacher: 47.3) while reducing network latency from 128.4 ms to 67.6 ms per frame. Latency is measured over network modules only and excludes simulator stepping, sensor acquisition, preprocessing, and I/O overhead; the resulting 14.8 FPS sustains a 10 Hz control cycle, meaning near-real-time operation throughout.