A Global–Local Differencing Network for Lightweight Point-Cloud Human Action Recognition
Fang Tan, Yupeng MaExisting methods for human action recognition from dynamic point clouds commonly rely on farthest point sampling and dense spatiotemporal neighborhood queries. The resulting local geometric computations are expensive, which complicates deployment in resource-constrained settings. This paper presents the Global–Local Differencing Network (GLD-Net), a lightweight framework for point-cloud sequence learning designed around feature extraction, motion representation, and temporal modeling. Feature extraction uses a two-branch architecture: the global branch encodes the complete point cloud in each frame to learn a holistic spatial representation, whereas the local branch uniformly divides the body along the vertical direction into several semantic regions, each processed by an independent network. Motion is represented directly by the distances from each point to its nearest neighbors in the preceding and subsequent frames, without requiring point correspondences. For temporal modeling, bidirectional differencing is applied to frame-level features to represent action changes explicitly. The method requires neither farthest point sampling nor complex spatiotemporal neighborhood searches. On MSR-Action3D, GLD-Net achieves 95.82% accuracy with 0.579 M parameters and 1.42 G operations. Compared with PvNeXt, a model of similar scale, GLD-Net improves accuracy by 1.05 percentage points while reducing the parameter count by 19.6%. The implementation code is publicly available.