ArticleBMC medical informatics and decision making2026
An integrated evaluation framework for synthetic clinical data in severely imbalanced settings: fidelity, privacy-risk profiling, and diagnostic utility.
Article in BMC medical informatics and decision making, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors.
Funding
Abstract
backgroundThe development of clinical artificial intelligence models is constrained by limited access to high-quality electronic health record data, a challenge that is particularly pronounced in rare and highly imbalanced clinical cohorts. Synthetic data generation has been proposed as a strategy to mitigate data-sharing barriers. However, an integrated evaluation framework that jointly examines distributional fidelity, diagnostic behavior, and privacy risk under such conditions remains lacking.
methodsWe developed an integrated evaluation framework to assess variational autoencoders (VAE) and conditional generative adversarial networks (CTGAN). The framework jointly characterizes distributional fidelity, privacy risk, and diagnostic behavior using structured electronic health record data from the KoGES cohort, with disease prevalence ranging from 0.8% to 7.5%, adopting a prevalence-aware approach in which evaluation metrics are stratified by disease-specific class frequency. To address the sensitivity of p-value-based tests in large samples, distance-based metrics, including Jensen-Shannon divergence and Wasserstein distance, were employed. Diagnostic behavior was evaluated using XGBoost, random forest, and logistic regression classifiers, with emphasis on minority-class-sensitive metrics such as recall and Macro-F1.
resultsIn multivariate structural analyses, the correlation similarity between empirical and synthetic data was 0.794 for VAE-generated data and 0.667 for CTGAN-generated data. Across diseases with moderate outcome prevalence, multivariate and stratified distributions exhibited numerical overlap between VAE-generated and empirical data. In diagnostic evaluations, classifiers trained on empirical data alone yielded zero recall for the rarest outcome (0.8% prevalence), whereas CTGAN-trained classifiers produced non-zero recall values at the cost of reduced overall accuracy. Across evaluated threat models, membership inference attack performance remained near the random-guessing reference (AUROC ≈ 0.500).
conclusionsThis study presents an integrated, prevalence-aware evaluation framework for synthetic clinical data that systematically identified failure modes undetectable by conventional single-metric approaches, including minority-class metric instability, heterogeneous tail-end privacy exposure, and qualitative divergence in membership score calibration. The evaluated generative models exhibited distinct trade-off profiles: VAE preserved multivariate structure while exhibiting lower proximity-based exposure under the evaluated threat model, whereas CTGAN achieved higher minority-class detection at the cost of structural fidelity. Supplementary augmentation experiments confirmed that increased synthetic data volume does not uniformly improve minority-class detection under extreme prevalence constraints. These findings demonstrate that numerical overlap between synthetic and empirical distributions does not guarantee clinical equivalence, underscoring the need for prevalence-stratified, multi-dimensional validation in structured EHR research.
trial registrationNot applicable.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.