DOI: 10.1177/20552076261467828 ISSN: 2055-2076

Validation of federated analytics across secure data environments: A comparative study of synthetic healthcare datasets using machine learning and general linear models

Suzy Gallier, Alexander Topham, James Hodson, David McNulty, Tom Charles Giles, Sam Cox, Mahalakshmi J. Chaganty, Lauren Cooper, Stephen Perks, Philip R. Quinlan, Elizabeth Sapey

Objectives

To enable research and innovation, most health data systems are moving away from a model of pooled data egress and instead highlight the benefits of federated analytics to support data analysis across different settings. This study aimed to assess whether the results of analyses using general linear models (GLMs) and machine learning (ML) models were altered depending on whether a federated or pooled data (“non-federated”) approach was taken.

Methods

A Conditional Transformation Generative Adversarial Network created two synthetic datasets (training set: N=10,000; test set: N=1,000), using data from 381 asthma patients. Synthetic data was either distributed across Trusted Research Environments (TREs) or pooled. GLMs (one-way analysis of variance) and ML models (gradient boosted decision trees) were produced, using federated and non-federated approaches. The consistency of predictions produced by the ML models and GLM results were compared.

Results

Federated GLMs were identical to those produced using a non-federated approach. However, ML models produced by federated and non-federated approaches, and using different data distributions between TREs, were non-identical. Despite this, applying the ML models to the test set returned similar predictive accuracies (area under the receiver operating characteristic curve: 0.663-0.669) and classification accuracies (84.7-90.4%) for federated and non-federated models.

Conclusions

GLMs return identical models for federated and pooled data approaches. Whilst federation of ML analysis results in non-identical models, these have comparable accuracy to traditional non-federated models. These results highlight the viability of federated approaches for reliable and accurate data analysis in sensitive domains.

More from our Archive