DOI: 10.3390/electronics15163609 ISSN: 2079-9292

A Generalized Slimmable INR Framework for Scalable Video Coding

Qingyu Mao, Jiacong Chen, Shuai Liu, Fanyang Meng, Yongsheng Liang, Youneng Bao

Implicit neural representations (INRs) encode video frames as network weights, offering a new compression paradigm. A persistent limitation is that existing INR codecs train one model per target bitrate, so multi-rate deployment needs separate runs, separate checkpoints, and model reloading, costs that grow with the number of rate points. We propose a Generalized Slimmable Framework that replaces standard layers with width-configurable counterparts in most INR decoders, letting a single checkpoint serve multiple bitrates through nested weight tensors without topology changes. To recover the quality lost in shared-weight training, we introduce Slimmable Conditional Decoder Modulation (SCDM), which blends slimmable expert convolutions via width- and frame-conditioned gating. For encoder-based backbones, encoder output caching separates encoder computation from multi-width training. Across four INR backbones on DAVIS and Bunny, the framework cuts multi-rate storage by about 2.3–2.5×, and SCDM recovers substantial quality at all rate points while surpassing independently trained fixed-width models at narrow widths.

More from our Archive