DOI: 10.3390/s26154854 ISSN: 1424-8220

BiasFormer: Structured Posterior-Guided Transformer for Temporal Action Segmentation

Dongyue Zhou, Zhihui Shi, Haixia Wang, Hanqing Yang, Mingyue Yang, Chongchong Yu

Transformer-based architectures have become a dominant paradigm for temporal action segmentation because of their ability to capture long-range temporal dependencies. However, content-driven self-attention does not explicitly model action durations or feasible action transitions, which can lead to short spurious fragments and ambiguous action boundaries. To address this limitation, we propose BiasFormer, a unified end-to-end framework that couples feature-level structural guidance with decoder-level structural constraints. Specifically, frame-level confidence biases derived from structured Markov posteriors guide reliability-aware feature reweighting in frame-to-action cross-attention. We further introduce the unified segmental Markov head, which formulates temporal action segmentation as a segmental conditional random field with explicit duration modeling and transition support constraints. The structured objective is jointly optimized with the backbone, enabling structural information to influence both feature interaction and decoding. Extensive experiments on Breakfast, GTEA, EgoProceL, and EPIC-KITCHENS demonstrate competitive performance and improvements over FACT on most segment-level metrics. These results support the effectiveness of coupling posterior-guided feature refinement with structured decoding for improving segment-level temporal coherence.

More from our Archive