Evidence map›Paper›PMID 39232851›Full record

ArticleBioData mining2024

Processing imbalanced medical data at the data level with assisted-reproduction data as an example.

Junliang Zhu, Shaowei Pu, Jiaji He, Dongchao Su, Weijie Cai, Xueying Xu, Hongbo Liu

Abstract read
In one paragraph

Article in BioData mining, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 11 papers.

0numbers the graph read from it
0cells of the map it votes in
11citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

11 citing papers in PubMed.

  1. Review
  2. Article
  3. Article
  4. Augmenting Electronic Health Records for Adverse Event Detection.medRxiv : the preprint server for health sciences · 2026
    Article
  5. Article
  6. An explainable machine learning approach to predicting carbapenem resistance inFrontiers in cellular and infection microbiology · 2026
    Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Junliang ZhuDepartment of Health Statistics, School of Public Health, China Medical University, Shenyang, 110122, PR China.
Shaowei PuDepartment of Health Statistics, School of Public Health, China Medical University, Shenyang, 110122, PR China.
Jiaji HeDepartment of Health Statistics, School of Public Health, China Medical University, Shenyang, 110122, PR China.
Dongchao SuDepartment of Health Statistics, School of Public Health, China Medical University, Shenyang, 110122, PR China.
Weijie CaiDepartment of Health Statistics, School of Public Health, China Medical University, Shenyang, 110122, PR China.
Xueying XuDepartment of Health Statistics, School of Public Health, China Medical University, Shenyang, 110122, PR China.
Hongbo LiuDepartment of Health Statistics, School of Public Health, China Medical University, Shenyang, 110122, PR China. hbliu@cmu.edu.cn.

Funding

The Science and Technology Planning Project of Liaoning Province 2021JH4/10200008The Science Research Project of Education Department of Liaoning Province LJKZ0765The Science Research Project of Shenyang City 23-506-3-01-21
6 · The paper itself

Abstract

objectiveData imbalance is a pervasive issue in medical data mining, often leading to biased and unreliable predictive models. This study aims to address the urgent need for effective strategies to mitigate the impact of data imbalance on classification models. We focus on quantifying the effects of different imbalance degrees and sample sizes on model performance, identifying optimal cut-off values, and evaluating the efficacy of various methods to enhance model accuracy in highly imbalanced and small sample size scenarios.

methodsWe collected medical records of patients receiving assisted reproductive treatment in a reproductive medicine center. Random forest was used to screen the key variables for the prediction target. Various datasets with different imbalance degrees and sample sizes were constructed to compare the classification performance of logistic regression models. Metrics such as AUC, G-mean, F1-Score, Accuracy, Recall, and Precision were used for evaluation. Four imbalance treatment methods (SMOTE, ADASYN, OSS, and CNN) were applied to datasets with low positive rates and small sample sizes to assess their effectiveness.

resultsThe logistic model's performance was low when the positive rate was below 10% but stabilized beyond this threshold. Similarly, sample sizes below 1200 yielded poor results, with improvement seen above this threshold. For robustness, the optimal cut-offs for positive rate and sample size were identified as 15% and 1500, respectively. SMOTE and ADASYN oversampling significantly improved classification performance in datasets with low positive rates and small sample sizes.

conclusionsThe study identifies a positive rate of 15% and a sample size of 1500 as optimal cut-offs for stable logistic model performance. For datasets with low positive rates and small sample sizes, SMOTE and ADASYN are recommended to improve balance and model accuracy.

Indexed as

Imbalanced dataImbalanced data processing methodImbalanced degreeLogistic modelSample size

Identifiers

PMID39232851
PMCPMC11373105

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.