Use of Federated Learning for validating and updating privacy-preserving decentralized multi-study prognostic models in Traumatic Brain Injury

Wait 5 sec.

Developing modern clinical prediction models (CPMs) and advanced analytics requires large datasets, often necessitating data from different studies. Privacy regulations may hinder data sharing, especially across countries. Decentralized federated data infrastructures, where data remain in their original location and analyses are run only in a shared, secure environment, may address these challenges. We implemented a privacy-preserving federated learning (FL) infrastructure and evaluated and updated the IMPACT prognostic models for traumatic brain injury (TBI) using 2 studies. A multi-continental federated infrastructure was established between 2 large-scale studies (TRACK-TBI from the United States and CENTER-TBI from Europe and Israel). Three IMPACT prognostic models for post-TBI 6-month mortality and unfavorable outcomes were evaluated, followed by model updates through 2 FL approaches trained across the TRACK-TBI and CENTER-TBI studies. Internal validation, external cross-validation, and sub-study validations were performed. CPMs were evaluated for discrimination and calibration. The federated cohort included 1616 participants (TRACK-TBI: n=441, CENTER-TBI: n=1175). Both FL performed well, with comparable coefficient estimates, AUCs (area under the receiver operating characteristics curve) between 0.77-0.88, and calibrated probabilities. Compared to the original IMPACT and single-study models, both federated models presented similar discrimination (AUC), were well-calibrated, were more efficient (higher precision), and reduced the impact of missing data in model estimation. FL is feasible for privacy-preserving development and evaluation of CPMs, and can enable validation and updating across large, virtually analyzed datasets while overcoming regulatory constraints on data combination. Federated infrastructures can facilitate global collaboration to advance data-hungry analytical methods, such as artificial intelligence.