SEFA: Semantic Embedding-Based Feature Augmentation of Biomedical Language-Model Embeddings Improves Interpretable Metabolomic Prediction of Lung Cancer
Jiawen Wu, Jean-François Haince, Rashid A. Bux, Guoyu Huang, Paramjit S. Tappia, Bram Ramjiawan, Maria VaidaBackground/Objectives: Feature engineering remains a major challenge in metabolomics-based prediction, particularly when rich biochemical knowledge is available but underutilized. Conventional metabolomics models rely primarily on measured variables and statistically driven feature selection, overlooking the molecular and pathway context encoded in curated metabolite knowledge bases. We propose SEFA (Semantic Embedding-based Feature Augmentation), a model-agnostic framework that integrates metabolite-level textual knowledge from the Human Metabolome Database (HMDB) into structured metabolomics modeling for lung-cancer prediction. Methods: SEFA encodes HMDB metabolite descriptions as 768-dimensional MedBERT vectors and projects measured metabolite concentrations into this semantic space via concentration-weighted aggregation, producing a 928-dimensional candidate feature matrix that concatenates 11 clinical variables, 149 metabolite concentrations, and 768 semantic projection features. Sparse L1-guided feature selection reduced this representation to 24 features (2.6% of candidates) within a leakage-free cross-validation pipeline. Six classifiers were evaluated on a lung-cancer plasma metabolomics cohort of 800 participants (586 cases, 214 controls) with a stratified 80:20 split, and a controlled ablation study compared the augmented representation with a 24-metabolite-only baseline. Pathway-enrichment analysis of the metabolite sets associated with the retained embedding dimensions was performed using MetaboAnalyst 5.0. Results: Logistic regression on the 24-feature SEFA representation achieved a test ROC-AUC of 0.969, a precision-recall AUC of 0.987, and an accuracy of 94.4%, competitive with less interpretable approaches. Under the same 24-feature budget, embedding augmentation improved test ROC-AUC by 0.008, precision-recall AUC by 0.004, and accuracy by 3.1 percentage points over the metabolite-only baseline. Five retained embedding dimensions mapped onto coherent metabolic themes—sphingolipids and acylcarnitines, carnitine and glutamine metabolism, one-carbon and nitrogen handling, purine catabolism, and oxidative stress markers—and their associated metabolite sets were enriched for arginine and proline metabolism and glycine, serine, and threonine metabolism, pathways with established roles in lung-cancer biology. Conclusions: SEFA demonstrates that semantic embeddings derived from biomedical language models can convert curated metabolite annotations into patient-level features that supply complementary predictive signal while preserving biological interpretability through pathway-level analysis. The present evidence is limited to a single region-specific cohort; external, cross-platform, and cross-disease evaluation is required before broader generalization or clinical application.