Benchmarking token mixers for efficient transformer-based learned video compression
Chun Zhang, Heming Sun, Jiro KattoLearned video compression has rapidly evolved, with recent approaches demonstrating potential in complex context modeling. However, these performance gains often come at the cost of significant complexity in both framework design and computation-heavy modules, obscuring the efficiency contribution of the core architectural components. This work utilizes the video compression transformer (VCT) as a controlled testbed to isolate and quantify the contribution of the transformer-based entropy model by its core context modeling schemes. The authors propose an abstracted transformer entropy model (ATEM) with a modular design and systematically evaluate six distinct token mixers – defined here as the core mechanisms responsible for feature propagation across tokens in the transformer architecture. These range from parameter-free pooling (Pooling-Mixer), hybrid convolution-attention blocks (Attn-Conv/Conv-Attn Hybrid Mixer), to improved sliding-window attention mechanisms (Swin-Mixer); each provides an insight into the compression task’s unique properties. Our optimal sliding window model achieved 14.86% reduction in BD-rate accompanied by a 17.6% decrease in computational cost, while the best performing CNN-Attention hybrid model yielded 16.64% BD-rate reduction with only a 1.73% increase in overhead compared to the VCT baseline. Through these experiments, this study establishes a clear lower bound and optimal trade-off points for context modeling in ATEM. Furthermore, this study provides critical insights into the redundancy of context module designs in video compression and establishes a new efficiency benchmark for future lightweight framework designs.