ICF-Fusion: Multimodal In-Cabin Sensor Fusion for Adaptive Restraint Systems
Victor Preu, Daniel Pauer, Roman Putter, Peter HeckerAdaptive restraint systems require specific occupant information, including head position, anthropometry, and safety-relevant posture states. Existing 3D human pose estimation benchmarks mostly report root-relative pose, while automotive in-cabin studies rarely evaluate these outputs across heterogeneous vehicle sensor sets. We present ICF-Fusion, a five-modality transformer fusion architecture, and evaluate it under leave-one-subject-out (LOSO) validation on the ICF-Body dataset, which includes synchronized near-infrared (NIR) camera, 60 GHz millimeter-wave (mmWave) radar, belt webbing extraction sensor (WES), seat configuration sensor (SCS), and ultra-wideband (UWB) recordings. The model localizes the head with a Mean Root Position Error (MRPE) of 6.10 cm and regresses anthropometry to mean absolute errors (MAE) of 5.36 cm for height, 3.50 cm for torso length, 1.78 cm for shoulder width, and 8.61 kg for weight. Feet-on-dashboard is detected on 9 of 10 evaluable folds without meaningful MRPE degradation. The full sensor fusion outperformed every single modality on all three tasks, but NIR alone nearly matched it for head localization and feet-on-dashboard detection. The fusion advantage was substantial only for the anthropometry estimation task.