Abstract PR019: Robust identification of recurrent gene expression programs across samples from the Single-cell Pediatric Cancer Atlas Portal
Allegra G. Hawkins, Joshua A. Shapiro, Stephanie J. Spielman, Jaclyn N. TaroniAbstract
With the rise of single-cell technologies, it has become increasingly apparent that pediatric tumors exhibit substantial transcriptomic heterogeneity. Many studies have even shown that specific tumor cell subpopulations and recurrent gene expression programs are associated with clinical outcome and may serve as therapeutic targets, revealing the importance of properly identifying these programs. Typical workflows for identifying recurrent gene expression programs, also referred to as metaprograms, use non-negative matrix factorization (NMF) on each individual sample and then combine these individual NMF programs into metaprograms using hierarchical clustering. However, this approach relies on both selecting the number of factors for NMF and choosing the appropriate number of clusters for hierarchical clustering. Because each dataset has its own unique characteristics, it is difficult to pre-determine the appropriate number of clusters or metaprograms. To address this problem, we curated a set of metrics to evaluate what constitutes a robust metaprogram: 1) non-redundancy, 2) sample diversity, and 3) biological interpretability. Here we present metafactory, a Nextflow workflow to identify metaprograms using a data-driven approach to select the optimal number of metaprograms (k) for each sample group. First, cNMF is run on each individual sample across a range of ranks. Sample-specific NMF programs are removed, and all remaining NMF programs across all samples are then clustered into metaprograms based on their Pearson correlation coefficient. The process of generating metaprograms is repeated across a wide range of k values, and a set of data-driven metrics are calculated for each value of k to assess the quality of resulting metaprograms. Each metric is used to rank the values of k, and the average rank is used to determine the optimal value of k and identify the final set of metaprograms for each sample group. This workflow can be applied to any set of single-cell RNA-sequencing datasets to identify recurrent gene expression programs. We applied metafactory to the Single-cell Pediatric Cancer Atlas (https://scpca.alexslemonade.org/), a data resource for uniformly processed single-cell and single-nuclei RNA sequencing data that contains data from over 700 samples across more than 50 cancer types. The use of metafactory provided a data-driven approach to identify recurrent gene expression programs across multiple disease types represented in the ScPCA Portal, accelerating the discovery of key programs that shape disease biology and may carry clinical significance.
Citation Format:
Allegra G. Hawkins, Joshua A. Shapiro, Stephanie J. Spielman, Jaclyn N. Taroni. Robust identification of recurrent gene expression programs across samples from the Single-cell Pediatric Cancer Atlas Portal [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Bridging Discovery and Clinical Impact in Pediatric Cancer; 2026 Sep 22-25; Philadelphia, PA. Philadelphia (PA): AACR; Cancer Res 2026;86(18_Suppl_1):Abstract nr PR019.