Abstract:A large amount of personal health data that is highly valuable to the scientific community is still not accessible or requires a lengthy request process due to privacy concerns and legal restrictions. As a solution, synthetic data has been studied and proposed to be a promising alternative to this issue. However, generating realistic and privacy-preserving synthetic personal health data retains challenges such as simulating the characteristics of the patients' data that are in the minority classes, capturing the relations among variables in imbalanced data and transferring them to the synthetic data, and preserving individual patients' privacy. In this paper, we propose a differentially private conditional Generative Adversarial Network model (DP-CGANS) consisting of data transformation, sampling, conditioning, and network training to generate realistic and privacy-preserving personal data. Our model distinguishes categorical and continuous variables and transforms them into latent space separately for better training performance. We tackle the unique challenges of generating synthetic patient data due to the special data characteristics of personal health data. For example, patients with a certain disease are typically the minority in the dataset and the relations among variables are crucial to be observed. Our model is structured with a conditional vector as an additional input to present the minority class in the imbalanced data and maximally capture the dependency between variables. Moreover, we inject statistical noise into the gradients in the networking training process of DP-CGANS to provide a differential privacy guarantee. We extensively evaluate our model with state-of-the-art generative models on personal socio-economic datasets and real-world personal health datasets in terms of statistical similarity, machine learning performance, and privacy measurement. We demonstrate that our model outperforms other comparable models, especially in capturing the dependence between variables. Finally, we present the balance between data utility and privacy in synthetic data generation considering the different data structures and characteristics of real-world personal health data such as imbalanced classes, abnormal distributions, and data sparsity.

On the Trade-Off between Fidelity, Utility and Privacy of Synthetic Patient Data

Fidelity and Privacy of Synthetic Medical Data

On the Fidelity-Privacy Tradeoff of Synthetic Cancer Registry Data

Generating high-fidelity synthetic patient data for assessing machine learning healthcare software

Privacy Risk Assessment for Synthetic Longitudinal Health Data

On Utility and Privacy in Synthetic Genomic Data

Harnessing the power of synthetic data in healthcare: innovation, application, and privacy

Synthetic data for privacy-preserving clinical risk prediction

Fake It Till You Make It: Guidelines for Effective Synthetic Data Generation

Does Differentially Private Synthetic Data Lead to Synthetic Discoveries?

Downstream Fairness Caveats with Synthetic Healthcare Data

Controllable Synthetic Clinical Note Generation with Privacy Guarantees

Synthetic Data: Revisiting the Privacy-Utility Trade-off

Analyzing Medical Research Results Based on Synthetic Data and Their Relation to Real Data Results: Systematic Comparison From Five Observational Studies

Patient-centric synthetic data generation, no reason to risk re-identification in biomedical data analysis

Utility Assessment of Synthetic Data Generation Methods

PrivSyn: Differentially Private Data Synthesis

A primer on synthetic health data

A Scoping Review of Privacy and Utility Metrics in Medical Synthetic Data

Generating synthetic personal health data using conditional generative adversarial networks combining with differential privacy