DOI: 10.1093/rap/rkag114 ISSN: 2514-1775

Synthetic Data in Rheumatology: A Systematic Literature Review

Vincenzo Venerito, Maria Morrone, Sergio Del Vescovo, Emre Bilgin, Latika Gupta, Giuseppe Lopalco, Florenzo Iannone

Abstract

Objective

Synthetic data are artificially generated records that reproduce the statistical properties of real patient data without corresponding to any real individual.

To systematically review the methods, validation frameworks, applications, and regulatory landscape of synthetic data generation in healthcare, with a specific focus on rheumatology.

Methods

We searched PubMed/MEDLINE, Embase, and arXiv (January 2016 to January 2026). From 1,099 database records and 7 expert-identified studies, 701 remained after removing 405 duplicates. Title and abstract screening excluded 353 records; full-text assessment of 348 articles excluded 120 (E1: n = 62; E2: n = 1; E3: n = 23; E4: n = 34), yielding 228 studies. Quality assessment of 13 rheumatology-specific studies using a 10-item checklist revealed 1 high-quality, 8 moderate-quality, and 4 low-quality studies, with systematic gaps in privacy evaluation (2/13) and reproducibility (1/13).

Results

The 228 included studies encompassed six generation method categories (generative adversarial networks, diffusion models, large language models, statistical methods, rule-based simulators, and hybrid approaches) and five thematic categories (digital twins, privacy-preserving methods, reviews, validation frameworks, and healthcare applications), detailed in Supplementary Methods. Validation remained inconsistent, with only a minority evaluating fidelity, utility, and privacy simultaneously. In rheumatology, 13 studies spanned rheumatoid arthritis (n = 3), osteoarthritis (n = 3), Sjogren’s syndrome (n = 2), systemic sclerosis (n = 2), systemic lupus erythematosus (n = 2), and ankylosing spondylitis (n = 1). Only 2 of 13 formally evaluated privacy, and only 1 shared code for reproducibility.

Conclusions

Synthetic data holds substantial promise for rheumatology, particularly for augmenting small cohorts in low-prevalence diseases, enabling privacy-preserving registry sharing, and simulating clinical trial populations. Realizing this potential requires standardized validation, greater attention to privacy and reproducibility, and regulatory clarity on synthetic health data.