Multimodal Data Fusion for a Self-Adaptive, Smart, Serious-Game Ecosystem Under Development
Xiya Tao, Peng Chen, Martina EckertThis article presents the implementation and technical feasibility evaluation of the multimodal sensing and feature-level fusion layer of BLEXER v3, a broader serious-game ecosystem under development for upper-limb rehabilitation. The implemented framework integrates Kinect-based motion tracking, Polar H10 and Bangle.js physiological sensing, wearable accelerometer data, and facial affective cues within a middleware-based architecture. Heterogeneous sensor streams are locally preprocessed, temporally aligned, and transformed into a common quality-aware multimodal feature representation containing motion, heart-rate and heart-rate-variability-related descriptors, affective information, availability indicators, signal-quality metadata, and freshness descriptors. The fused representation is additionally mapped, using predefined rules, to heuristic operational descriptors, including low demand, moderate stable, active engagement, physical load, affective activation, high demand, and uncertain. These descriptors are not intended as clinical diagnoses, independently validated user states, or final adaptation decisions. An exploratory K-means analysis of 12,675 complete multimodal windows reveals partial correspondence between the data-driven cluster structure and the predefined operational descriptors. Some descriptors show comparatively concentrated cluster patterns, whereas others exhibit overlap and internal heterogeneity. The results demonstrate the technical feasibility of generating structured multimodal representations that can provide input for subsequent context-aware reasoning. Independent validation of the operational descriptors, completion and evaluation of the whole system, and clinical validation with rehabilitation patients remain future work.