DOI: 10.1158/1538-7445.pancreatic26-a104 ISSN: 0008-5472

Abstract A104: Transfer learning for Pancreatic ductal adenocarcinoma (PDAC) subtyping using extracellular vesicle RNA (evRNA)

Dinelka Nanayakkara, Anirban Maitra, Vince Bernard-Pagan, Xianlu L. Peng, Jen Jen Yeh, Naim U. Rashid

Abstract

Identifying the molecular subtype of pancreatic ductal adenocarcinoma (PDAC) hinges on tissue biopsies requiring invasive surgical procedures. The PurIST classifier identifies two molecular subtypes: basal-like and classical. PDAC is commonly diagnosed at an advanced stage, making invasive measures risky. Liquid biopsies such as extracellular vesicle RNA (evRNA) have recently become popular to overcome such problems. However, alongside the value it brings, evRNA has its limitations: high gene expression dropout. In real data (n=39), certain genes are not expressed in at least 70% of samples. This high dropout/missingness is not random, but is dependent on the gene expression value itself. In contrast, bulkRNA gene expression, obtained via tissue biopsies, does not exhibit this limitation. We introduce the geneBTLM classifier, trained on bulkRNA and evRNA, with the goal of predicting PDAC subtype on unseen evRNA expression. geneBTLM is built using Bayesian factor modeling, which estimates gene loadings and latent factors for both bulkRNA and evRNA expression. Paired samples are those that have both bulk and evRNA expression data. Latent factors for paired samples remain equal between the two sources. Gene loadings in evRNA are a scaled version of those in bulkRNA. Missingness and probability of subtype are modeled using logistic regression. The subtype classifier assumes the outcome to depend on the estimated latent factors. Subtype prediction and corresponding credible intervals on a new evRNA sample are obtained using the posterior predictive distribution. Gibbs sampling is used for parameter estimation. Given that our model can’t be tested on samples with no true subtype, only samples that had bulkRNA were used for the train/test steps. The number of latent factors was determined using either 5-fold or leave-one-out cross-validation (LOOCV), depending on the sample size of our training data. Our main simulations consisted of 700 tissue biopsy only samples, 200 paired samples, and 100 liquid biopsy only samples. Real data consisted of 1300 publicly available bulkRNA samples in which the labels were generated using the PurIST classifier, and 39 paired evRNA samples. 5 of these 39 paired evRNA samples were held out for testing. In simulations, the 95% Bayesian posterior credible intervals of the Area under the Curve (AUC) were: (1) (0.78, 0.84) with no missingness, (2) (0.60, 0.76) with simulated missingness, and (3) (0.46, 0.69), when missingness levels are scaled higher than that in (2). Further testing on other simulation scenarios is ongoing. Of the 5 held-out samples in real data, the 95% Bayesian posterior credible interval for the predicted probability of being basal-like in the true basal-like sample was (0.35, 0.65), and the mean across the 4 true classical samples was (0.10, 0.32). While not yet clinically translatable, geneBTLM takes a step forward toward a PDAC subtype classifier using a blood draw, and we hope these results will motivate researchers to generate more evRNA-seq data to validate and further extend this method.

Citation Format:

Dinelka Nanayakkara, Anirban Maitra, Vince Bernard-Pagan, Xianlu L. Peng, Jen Jen Yeh, Naim U. Rashid. Transfer learning for Pancreatic ductal adenocarcinoma (PDAC) subtyping using extracellular vesicle RNA (evRNA) [abstract]. In: Proceedings of the AACR Conference on Pancreatic Cancer: New Frontiers in Biology and Therapeutic Development; 2026 Sep 25-28; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(18_Suppl_2):Abstract nr A104.