ArticleBioData mining2024
Processing imbalanced medical data at the data level with assisted-reproduction data as an example.
Article in BioData mining, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 11 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
11 citing papers in PubMed.
- Artificial Intelligence-Driven Reproductive Bioengineering: Integrating Fertility Diagnostics, Organ-on-Chip Systems, Cryobiology and Epigenetic Safety for Precision Reproductive Medicine.Bioengineering (Basel, Switzerland) · 2026Review
- Prediction of cancer-associated thrombosis by machine learning: results from the Vienna Cancer and Thrombosis Study.ESMO open · 2026Article
- A TyG-UHR-based machine learning model for screening lean MAFLD: development and external validation.Biomedical engineering online · 2026Article
- Augmenting Electronic Health Records for Adverse Event Detection.medRxiv : the preprint server for health sciences · 2026Article
- Effect of air pollution on clinical pregnancy outcomes following intrauterine insemination: a machine learning-based analysis.Frontiers in public health · 2026Article
- An explainable machine learning approach to predicting carbapenem resistance inFrontiers in cellular and infection microbiology · 2026Article
- Evaluation of Machine Learning Model Performance in Diabetic Foot Ulcer: Retrospective Cohort Study.JMIR medical informatics · 2025Article
- Development of a single-center predictive model for conventional in vitro fertilization outcomes excluding total fertilization failure: implications for protocol selection.Journal of ovarian research · 2025Article
- Construction and validation of machine learning models for predicting lymph node metastasis in cutaneous malignant melanoma: a large population-based study.Translational cancer research · 2025Article
- Evaluating machine learning models for stroke prediction based on clinical variables.Frontiers in neurology · 2025Article
- Predicting no-shows at outpatient appointments in internal medicine using machine learning models.PeerJ. Computer science · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
7 authors.
Funding
Abstract
objectiveData imbalance is a pervasive issue in medical data mining, often leading to biased and unreliable predictive models. This study aims to address the urgent need for effective strategies to mitigate the impact of data imbalance on classification models. We focus on quantifying the effects of different imbalance degrees and sample sizes on model performance, identifying optimal cut-off values, and evaluating the efficacy of various methods to enhance model accuracy in highly imbalanced and small sample size scenarios.
methodsWe collected medical records of patients receiving assisted reproductive treatment in a reproductive medicine center. Random forest was used to screen the key variables for the prediction target. Various datasets with different imbalance degrees and sample sizes were constructed to compare the classification performance of logistic regression models. Metrics such as AUC, G-mean, F1-Score, Accuracy, Recall, and Precision were used for evaluation. Four imbalance treatment methods (SMOTE, ADASYN, OSS, and CNN) were applied to datasets with low positive rates and small sample sizes to assess their effectiveness.
resultsThe logistic model's performance was low when the positive rate was below 10% but stabilized beyond this threshold. Similarly, sample sizes below 1200 yielded poor results, with improvement seen above this threshold. For robustness, the optimal cut-offs for positive rate and sample size were identified as 15% and 1500, respectively. SMOTE and ADASYN oversampling significantly improved classification performance in datasets with low positive rates and small sample sizes.
conclusionsThe study identifies a positive rate of 15% and a sample size of 1500 as optimal cut-offs for stable logistic model performance. For datasets with low positive rates and small sample sizes, SMOTE and ADASYN are recommended to improve balance and model accuracy.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.