A Generalized Slimmable INR Framework for Scalable Video Coding
Qingyu Mao, Jiacong Chen, Shuai Liu, Fanyang Meng, Yongsheng Liang, Youneng BaoImplicit neural representations (INRs) encode video frames as network weights, offering a new compression paradigm. A persistent limitation is that existing INR codecs train one model per target bitrate, so multi-rate deployment needs separate runs, separate checkpoints, and model reloading, costs that grow with the number of rate points. We propose a Generalized Slimmable Framework that replaces standard layers with width-configurable counterparts in most INR decoders, letting a single checkpoint serve multiple bitrates through nested weight tensors without topology changes. To recover the quality lost in shared-weight training, we introduce Slimmable Conditional Decoder Modulation (SCDM), which blends slimmable expert convolutions via width- and frame-conditioned gating. For encoder-based backbones, encoder output caching separates encoder computation from multi-width training. Across four INR backbones on DAVIS and Bunny, the framework cuts multi-rate storage by about 2.3–2.5×, and SCDM recovers substantial quality at all rate points while surpassing independently trained fixed-width models at narrow widths.