DOI: 10.1145/3837863 ISSN: 1551-6857

Context-aware and View-consistent Learning for Multi-view Action Recognition

Trung Thanh Nguyen, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide

The increasing deployment of multi-sensor systems in smart homes, surveillance, and assistive robotics has intensified research on multi-view action recognition. While existing research is effective under fully overlapping sensor configurations, real-world systems typically exhibit partially overlapping views, where actions are visible only in a subset of sensors. This non-uniform visibility, coupled with severe occlusions and inconsistent view coverage across sensors, poses a major challenge for reliable action recognition. Moreover, most current methods focus on implicit feature fusion and overlook explicit reasoning about contextual cues and prediction-level consistency across views, leading to degraded performance under occlusion and misalignment. To address these challenges, we propose Co ntext-aware and Vi ew-consistent learning ( CoVi ), a unified method that enhances multi-view action recognition by jointly modeling contextual dependencies and cross-view coherence. CoVi introduces two plug-and-play modules: Context-aware and View-aware modules. The former adaptively integrates local human-centric and global scene-level information through gated fusion, enabling the model to emphasize informative regions and suppress background noise for fine-grained contextual reasoning. The latter enforces prediction alignment across co-visible views using confidence-weighted Jensen–Shannon divergence, ensuring consistent learning under view disparity and partial visibility. Extensive experiments conducted on two real-world home environments from the MultiSensor-Home dataset and the MM-Office dataset demonstrate that CoVi consistently outperforms state-of-the-art baselines in both uni-modal and multi-modal settings, validating its effectiveness and generalizability across diverse sensor configurations and scene types.

More from our Archive