DOI: 10.4258/hir.2026.32.3.285 ISSN: 2093-369X

Machine Learning–Based Classification of Active and Latent Phases of Inherited Retinal Dystrophies Using Synthetic Proteomic Data: A Pathway-Based Application Exercise

Alessandro Macchia, Maria Chiara Medori, Ahmad Jainul Abidin, Gabriele Bonetti, Kristjana Dhuli, Luca Ferrari, Sara Feizyab, Jan Miertuš, Benedetto Falsini, Giorgio Placidi, Pietro Chiurazzi, Ornela Gordani, Xhilda Dhamo, Eglantina Kalluci, Dominika Vešelényiová, Stanislav Miertuš, Iveta Dirgová Luptáková, Jiří Pospíchal, Matteo Gregorini, Matteo Bertelli

Objectives: Inherited retinal dystrophies are characterized by high genetic and phenotypic heterogeneity, and their clinical progression may alternate between latent and active phases. Identifying the onset of the active phase may support earlier intervention for inflammatory retinal degeneration. Plasma proteomics has shown potential for characterizing predictive biomarkers in retinal diseases, but its application remains experimental. This study aimed to develop a methodological simulation exercise to evaluate the performance of machine learning (ML) models in distinguishing active and latent phases using an artificially generated dataset.Methods: An artificial dataset of 500 samples was created, and plasma proteomic profiles were generated for each sample using arbitrary values. Sample classification was based on a pathway activation score. Four ML models were tested: support vector machine, random forest, logistic regression, and extreme gradient boosting. Each model was trained across a range of hyperparameters.Results: Logistic regression achieved the best performance, with an accuracy of 0.73, precision of 0.73, and F1-score of 0.73.Conclusions: This simulation study showed that synthetic proteomic datasets can be used to evaluate ML approaches for distinguishing active and latent phases of retinal dystrophies when real data are scarce. Synthetic data can support the creation of targeted datasets for proteins associated with retinal dystrophies, helping to address the limited availability of suitable open-source data.

More from our Archive