ArticleBMC public health2025
Interpretable machine learning method to predict the risk of pre-diabetes using a national-wide cross-sectional data: evidence from CHNS.
Article in BMC public health, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
10 citing papers in PubMed.
- Development and Validation of an Explainable Machine Learning Model to Assess the Prevalence Probability of Gastrointestinal Heat Retention Syndrome in Children: Cross-Sectional Study.Journal of medical Internet research · 2026Article
- Development of a biochemical model for identifying prediabetes in Chinese adults: a case-control study.BMC endocrine disorders · 2026Article
- Comparison of the predictive power of novel insulin resistance, lipid metabolism and obesity indices for the risk of developing dyslipidaemia in adults.BMC endocrine disorders · 2026Article
- Artificial intelligence in prediabetes care: applications in screening, risk prediction, and lifestyle intervention.Frontiers in endocrinology · 2026Review
- From raw data to actionable insights: preprocessing real-world data for machine learning in diabetes care.Frontiers in digital health · 2026Article
- Artificial intelligence model as a tool to predict prediabetes.Scientific reports · 2025Article
- Bioinformatics mining and experimental validation of prognostic biomarkers in colorectal cancer.Discover oncology · 2025Article
- Applications of Artificial Intelligence and Machine Learning in Prediabetes: A Scoping Review.Journal of diabetes science and technology · 2025Review
- Machine Learning-Based Classification of Diabetes: Model Accuracy, Feature Importance, and Clinical Implications.Acta informatica medica : AIM : journal of the Society for Medical Informatics of Bosnia & Herzegovina : casopis Drustva za medicinsku informatiku BiH · 2025Article
- A machine learning model for predicting obesity risk in patients with diabetes mellitus: analysis of NHANES 2007-2018.Frontiers in public health · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
10 authors.
Funding
Abstract
objectiveThe incidence of Type 2 Diabetes Mellitus (T2DM) continues to rise steadily, significantly impacting human health. Early prediction of pre-diabetic risks has emerged as a crucial public health concern in recent years. Machine learning methods have proven effective in enhancing prediction accuracy. However, existing approaches may lack interpretability regarding underlying mechanisms. Therefore, we aim to employ an interpretable machine learning approach utilizing nationwide cross-sectional data to predict pre-diabetic risk and quantify the impact of potential risks.
methodsThe LASSO regression algorithm was used to conduct feature selection from 30 factors, ultimately identifying nine non-zero coefficient features associated with pre-diabetes, including age, TG, TC, BMI, Apolipoprotein B, TP, leukocyte count, HDL-C, and hypertension. Various machine learning algorithms, including Extreme Gradient Boosting (XGBoost), Random Forest (RF), Support Vector Machine (SVM), Naive Bayes (NB), Artificial Neural Networks (ANNs), Decision Trees (DT), and Logistic Regression (LR), were employed to compare predictive performance. Employing an interpretable machine learning approach, we aimed to enhance the accuracy of pre-diabetes risk prediction and quantify the impact and significance of potential risks on pre-diabetes.
resultsFrom the China Health and Nutrition Survey (CHNS) data, a cohort of 8,277 individuals was selected, exhibiting a disease prevalence of 7.13%. The XGBoost model demonstrated superior performance with an AUC value of 0.939, surpassing RF, SVM, DT, ANNs, Naive Bayes, and LR models. Additionally, Shapley Additive Explanation (SHAP) analysis indicated that age, BMI, TC, ApoB, TG, hypertension, TP, HDL-C, and WBC may serve as risk factors for pre-diabetes.
conclusionThe constructed model comprises nine easily accessible predictive factors, which prove highly effective in forecasting the risk of pre-diabetes. Concurrently, we have quantified the specific impact of each predictive factor on the risk and ranked them based on their influence. This result may serve as a convenient tool for early identification of individuals at high risk of pre-diabetes, providing effective guidance for preventing the progression of pre-diabetes to T2DM.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.