Sequencing Saturation Does Not Uniquely Determine Molecular Recovery in UMI Transcriptomics
Gavin W Wilson, Sangeetha N Kalimuthu, Jonathan C YeungAbstract
Motivation
Accurate sequencing depth planning for UMI-based transcriptomic experiments currently relies on heuristic metrics such as reads per cell and sequencing saturation. However, sequencing saturation cannot uniquely determine molecular recovery because the relationship between these quantities depends on amplification heterogeneity.
Results
Here we present NB-Lib, a modeling framework based on a zero-truncated negative binomial representation of reads per molecule that jointly estimates amplification heterogeneity and library complexity from transcriptomic libraries. Across 150 single-cell and spatial transcriptomic datasets, NB-Lib accurately reconstructs sequencing saturation curves and predicts sequencing depth requirements from shallow pilot sequencing experiments. We show that amplification heterogeneity and library complexity define a compact parameterization linking sequencing depth, saturation, and molecular recovery across transcriptomic platforms and explain substantial variation in sequencing cost and recovery between samples and technologies. Finally, we demonstrate that molecular recovery directly determines the reproducibility of low-abundance gene detection. Together, these results establish a unified framework for interpreting sequencing saturation, molecular recovery, and sequencing efficiency across UMI-based transcriptomic technologies.
Availability
The scdepth package is available at https://github.com/gwlab-ca/scdepth . Code needed to reproduce the analyses presented in this work is available at https://github.com/gwlab-ca/scdepth_manuscript . A selection of raw data and all the post-processed data is available at https://doi.org/10.5281/zenodo.15518941.
Supplementary information
Supplementary data are available at Bioinformatics online.