Benchmarking privacy and utility in synthetic tabular cohorts for Alzheimer's disease research
Filip Winzell, Ida Arvidsson, Niels Christian Overgaard, Anders Heyden, Kalle Åström, Linda Karlsson, Jacob W. Vogel, Oskar Hansson, Niklas Mattsson‐Carlgren,Abstract
INTRODUCTION
The scarcity of large, clinically relevant cohorts is becoming a bottleneck in Alzheimer's disease (AD) research, as their sensitive nature makes open data sharing difficult. Privacy‐preserving synthetic datasets generated with machine learning may help address this challenge.
METHODS
We compared five frameworks for generating synthetic tabular data from the Alzheimer's Disease Neuroimaging Initiative and Anti‐Amyloid Treatment in Asymptomatic Alzheimer's Disease cohorts, with a set of empirical privacy and utility metrics. Two of the methods, DataSynthesizer and TableDiffusion, provide ‐differential privacy guarantees.
RESULTS
Methods with differential privacy achieved high privacy ratings but low levels of utility. Deep learning methods like Tabular Prior‐data Fitted Network (TabPFN) and Conditional Generative Adversarial Network (CTGAN) also showed high privacy with limited utility. In contrast, non‐private DataSynthesizer and Synthpop offered higher utility at a cost of lower privacy.
DISCUSSION
The evaluated methods demonstrated a clear trade‐off between privacy and utility. High privacy was generally associated with insufficient utility, highlighting the need for further research into synthetic data generation for AD.