Evidence map›Paper›PMID 41639760›Full record

ArticleBMC medical research methodology2026

Machine learning performance for a small dataset: random oversampling improves data imbalances and fairness.

Lin Wang, Elliott Shi, Brett Meyers, Pavlos Vlachos, James Tcheng, Scott Denardo

Abstract read
In one paragraph

Article in BMC medical research methodology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Lin WangDepartment of Statistics, Purdue University, West Lafayette, IN, US.
Elliott ShiDepartment of Statistics, Purdue University, West Lafayette, IN, US.
Brett MeyersSchool of Mechanical Engineering, Purdue University, West Lafayette, IN, US.
Pavlos VlachosSchool of Mechanical Engineering, Purdue University, West Lafayette, IN, US.
James TchengDuke University Medical Center, Durham, NC, US.
Scott DenardoDuke University Medical Center, Durham, NC, US. scott.denardo@duke.edu.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundPercutaneous coronary intervention (PCI) can be complicated by major adverse cardiovascular events (MACE; death, myocardial infarction [MI], and target vessel revascularization). Statistical models facilitated by machine learning (ML) can improve prediction of MACE compared with conventional models. However, two key challenges impair ML performance: (1) class imbalance in outcome distributions; and (2) bias affecting fairness across social subgroups. These challenges are amplified in small datasets. Additionally, no published ML model specifically addresses the influence of social determinants of health (SDoH) on PCI outcomes. We hypothesized that random oversampling would improve sensitivity and fairness when assessing the effect of SDoH on select PCI-associated MACE in a small dataset.

methodsWe employed multivariable logistic regression to predict 180-day MACE and new MI following urgent PCI in a small dataset (N = 481). Three SDoH were pre-specified variables: race, marital status, and socioeconomic status at extremes of the spectrum (uninsured and Medicaid versus private insurance). Random oversampling was applied to the SDoH and outcomes while maintaining a consistent ratio between positive and negative outcomes.

resultsIn the imbalanced dataset, the logistic regression revealed no significant association between SDoH and 180-day MACE or new MI (minimum P-value = 0.47). Although sensitivity for event detection was low (≤ 0.26), other metrics-including positive predictive value (PPV)-were commendable (≥ 0.80). ML classifiers trained on oversampled data for race and marital status showed increased sensitivity but decreased PPV as the ratio of adverse-to-favorable outcome increased. Other performance metrics remained stable. Equalized odds disparity decreased with increasing oversampling ratio, indicating improved fairness. However, for socioeconomic-extremes status, the low numbers of adverse events and small sub-group size limited the effectiveness of oversampling. The optimal oversampling ratio cut-point across all outcomes and SDoH was 0.30-0.40 (sensitivity-0.50; PPV-0.52).

conclusionsIn small, imbalanced datasets, random oversampling can improve sensitivity and fairness of ML classifiers when evaluating PCI outcomes across race and marital status. However, this improvement comes at the expense of decreased PPV. The decrease in equalized odds disparity indicates improved balance in the performance of the classifiers across these two subgroups.

Indexed as

Machine LearningMyocardial InfarctionPercutaneous Coronary InterventionClassification AlgorithmsFemaleHumansLogistic ModelsPrediction AlgorithmsPredictive Learning ModelsSocial Determinants of HealthFairnessMachine learningMajor adverse cardiovascular eventsOversamplingPercutaneous coronary interventionSocial determinants of health

Identifiers

PMID41639760
PMCPMC13032653

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.