Evidence map›Paper›PMID 42351098›Full record

ArticleBMC medical informatics and decision making2026

An integrated evaluation framework for synthetic clinical data in severely imbalanced settings: fidelity, privacy-risk profiling, and diagnostic utility.

Youngtae Kim, Jungwoo Lee, Sang Baek Koh, Kyu Hee Lee

Abstract read
In one paragraph

Article in BMC medical informatics and decision making, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Youngtae KimArtificial Intelligence Big Data Medical Center, Yonsei University Wonju College of Medicine, Wonju, 26417, Korea.
Jungwoo LeeArtificial Intelligence Big Data Medical Center, Yonsei University Wonju College of Medicine, Wonju, 26417, Korea.
Sang Baek KohDepartment of Preventive Medicine, Yonsei University Wonju College of Medicine, Wonju, 26417, Korea.
Kyu Hee LeeDepartment of Precision Medicine, Yonsei University Wonju College of Medicine, Wonju, 26417, Korea. powerpc@yonsei.ac.kr.

Funding

Korea Evaluation Institute of Industrial Technology (KEIT) RS-2025-13882968National Research Foundation of Korea (NRF) RS-2025-24683718
6 · The paper itself

Abstract

backgroundThe development of clinical artificial intelligence models is constrained by limited access to high-quality electronic health record data, a challenge that is particularly pronounced in rare and highly imbalanced clinical cohorts. Synthetic data generation has been proposed as a strategy to mitigate data-sharing barriers. However, an integrated evaluation framework that jointly examines distributional fidelity, diagnostic behavior, and privacy risk under such conditions remains lacking.

methodsWe developed an integrated evaluation framework to assess variational autoencoders (VAE) and conditional generative adversarial networks (CTGAN). The framework jointly characterizes distributional fidelity, privacy risk, and diagnostic behavior using structured electronic health record data from the KoGES cohort, with disease prevalence ranging from 0.8% to 7.5%, adopting a prevalence-aware approach in which evaluation metrics are stratified by disease-specific class frequency. To address the sensitivity of p-value-based tests in large samples, distance-based metrics, including Jensen-Shannon divergence and Wasserstein distance, were employed. Diagnostic behavior was evaluated using XGBoost, random forest, and logistic regression classifiers, with emphasis on minority-class-sensitive metrics such as recall and Macro-F1.

resultsIn multivariate structural analyses, the correlation similarity between empirical and synthetic data was 0.794 for VAE-generated data and 0.667 for CTGAN-generated data. Across diseases with moderate outcome prevalence, multivariate and stratified distributions exhibited numerical overlap between VAE-generated and empirical data. In diagnostic evaluations, classifiers trained on empirical data alone yielded zero recall for the rarest outcome (0.8% prevalence), whereas CTGAN-trained classifiers produced non-zero recall values at the cost of reduced overall accuracy. Across evaluated threat models, membership inference attack performance remained near the random-guessing reference (AUROC ≈ 0.500).

conclusionsThis study presents an integrated, prevalence-aware evaluation framework for synthetic clinical data that systematically identified failure modes undetectable by conventional single-metric approaches, including minority-class metric instability, heterogeneous tail-end privacy exposure, and qualitative divergence in membership score calibration. The evaluated generative models exhibited distinct trade-off profiles: VAE preserved multivariate structure while exhibiting lower proximity-based exposure under the evaluated threat model, whereas CTGAN achieved higher minority-class detection at the cost of structural fidelity. Supplementary augmentation experiments confirmed that increased synthetic data volume does not uniformly improve minority-class detection under extreme prevalence constraints. These findings demonstrate that numerical overlap between synthetic and empirical distributions does not guarantee clinical equivalence, underscoring the need for prevalence-stratified, multi-dimensional validation in structured EHR research.

trial registrationNot applicable.

Indexed as

Electronic Health RecordsPrivacyAutoencoderBoosting Machine Learning AlgorithmsGenerative Adversarial NetworksGenerative Artificial IntelligenceHumansClass imbalanceConditional generative adversarial networkMembership inferencePrivacy-preserving machine learningSynthetic dataVariational autoencoder

Identifiers

PMID42351098
PMCPMC13560368

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.