DOI: 10.3390/s26196202 ISSN: 1424-8220

Driving-Context Classification from Wearable Physiological and Motion Signals During Real-World Driving

Poh Ping Em, Tai Shie Teoh

Wearable physiological sensing is widely used in driving research to infer drowsiness and stress, but whether driving context itself is reflected in wearable signals, independent of any drowsiness label, is rarely examined directly; any such signal reflects the driver’s physiological response to context, not the road environment itself. We analyzed 1536 non-overlapping 60-s windows from 18 driving sessions completed by 11 drivers wearing an Empatica EmbracePlus across three fixed real-world routes (Rural, Highway, Urban; round-trip distances 22.4–47.1 km). Because these windows are correlated pseudo-replicates of the 18-session experimental unit rather than independent observations, we report both the window-level mixed-effects comparison a naive analysis would present as primary (27 of 32 features significant after false-discovery-rate correction) and a corrected session-level analysis with driver-clustered standard errors, covariate adjustment for trip duration, sleep, pre-drive sleepiness, and route order. Using a single omnibus test per feature, only 2 of 32 features (skin conductance response count and accelerometer movement count) remain significant once the statistical unit matches the experimental unit and pairwise multiplicity is properly controlled. A random forest evaluated with leave-one-driver-out cross-validation achieved only 34.9% accuracy (macro-F1 = 0.33; 95% CI [20.0%, 49.9%]) against a 27.0–32.7% baseline, versus 90.6% accuracy (95% CI [89.1%, 92.0%]) under a naive ungrouped cross-validation that leaks driver identity across folds; the 55.6-point gap is clearly distinguishable from zero (95% CI [42.0, 69.2]). A four-way modality ablation (physiology, accelerometry, temperature, all combined) found no dominant sensor channel and the combined-channel model did not outperform accelerometry; these results are consistent with a diffuse cross-modal signature rather than a physiology-specific one. Driving context is statistically distinguishable from wrist-worn wearable signals once analyzed at the correct statistical unit, but the signal is diffuse across sensor channels, and person-independent classification remains weak at this sample size; naive cross-validation and uncorrected pairwise-multiplicity designs common in this literature can substantially overstate both classification and statistical testing results.