DOI: 10.3390/app16157797 ISSN: 2076-3417

Machine Learning Models and Gene Expression Profiles in Epileptic Samples: A Narrative Review

Claudia Cava, Soudabeh Sabetian, Maria Pagoni, Isabella Castiglioni

Epilepsy is a global public health problem with increasing prevalence. However, the quality and accuracy of diagnostic models in clinical practice remain limited. In recent years, machine learning (ML) models applied to transcriptomic data have been explored as potential tools for improving diagnostic accuracy. This narrative review aimed to evaluate studies that applied ML algorithms to gene expression data, including both mRNA and miRNA, for the diagnosis of epilepsy. PubMed and Scopus were searched for studies published between 2020 and June 2026. Eligible studies included human epilepsy samples, transcriptomic data was used to train ML classifiers, and only articles reporting AUC were included. A total of 17 studies were included. To provide a clinically meaningful synthesis, studies were grouped according to the biological source of the biomarkers into peripheral blood-based biomarkers and brain tissue-derived biomarkers. Peripheral blood studies investigated whole-blood transcriptomic signatures, peripheral blood mononuclear cells, circulating miRNAs, or extracellular vesicle-derived miRNAs, whereas brain tissue studies mainly analyzed hippocampal or cortical samples. Across studies, commonly used models included random forests, support vector machines, artificial neural networks, logistic regression, and penalized regression methods. Reported diagnostic performance was frequently high, but study designs were heterogeneous, sample sizes were often small, and validation strategies were limited. Studies with external validation generally showed more conservative and likely more realistic performance estimates than those relying on internal validation only. Overall, gene expression-based machine learning models show promising exploratory potential for epilepsy biomarker discovery and diagnostic stratification. However, the current evidence does not yet support clinical implementation because most studies remain limited by small sample sizes, heterogeneous preprocessing strategies, limited independent external validation, and potential risk of overfitting. Future studies should prioritize prospective cohorts, predefined analysis plans, harmonized preprocessing, nested cross-validation, external multicenter testing, and transparent reporting according to TRIPOD-AI/PROBAST-AI principles.

More from our Archive