Evidence map›Paper›PMID 40140819›Full record

ArticleBMC public health2025

Interpretable machine learning method to predict the risk of pre-diabetes using a national-wide cross-sectional data: evidence from CHNS.

Xiaolong Li, Fan Ding, Lu Zhang, Shi Zhao, Zengyun Hu, Zhanbing Ma, Feng Li, Yuhong Zhang, Yi Zhao, Yu Zhao

Abstract read
In one paragraph

Article in BMC public health, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers.

0numbers the graph read from it
0cells of the map it votes in
10citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

10 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Review
  5. Article
  6. Article
  7. Article
  8. Review
  9. Machine Learning-Based Classification of Diabetes: Model Accuracy, Feature Importance, and Clinical Implications.Acta informatica medica : AIM : journal of the Society for Medical Informatics of Bosnia & Herzegovina : casopis Drustva za medicinsku informatiku BiH · 2025
    Article
  10. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Xiaolong Li *School of Public Health, Ningxia Medical University, Yinchuan Ningxia, 750004, China.
Fan Ding *School of Public Health, Ningxia Medical University, Yinchuan Ningxia, 750004, China.
Lu ZhangSchool of Public Health, Ningxia Medical University, Yinchuan Ningxia, 750004, China.
Shi ZhaoSchool of Public Health, Tianjin Medical University, Tianjin, 300070, China.
Zengyun HuSchool of Public Health, Shanghai Jiao Tong University, Shanghai, 200025, China.
Zhanbing MaSchool of Basic Medicine, Ningxia Medical University, Yinchuan Ningxia, 750004, China.
Feng LiDepartment of Laboratory Medicine, General Hospital of Ningxia Medical University, Yinchuan Ningxia, 750004, China.
Yuhong ZhangSchool of Public Health, Ningxia Medical University, Yinchuan Ningxia, 750004, China.
Yi ZhaoSchool of Public Health, Ningxia Medical University, Yinchuan Ningxia, 750004, China. zhaoyi@nxmu.edu.cn.
Yu ZhaoSchool of Public Health, Ningxia Medical University, Yinchuan Ningxia, 750004, China. zhaoyu@nxmu.edu.cn.

Funding

Major Science and Technology Projects and Achievements "Open competition for selecting the best candidates" of Ningxia Medical University XJKF230204National Natural Science Foundation of China 12061058Western Light Talent Training Program of the Chinese Academy of Sciences XAB2022YW20
6 · The paper itself

Abstract

objectiveThe incidence of Type 2 Diabetes Mellitus (T2DM) continues to rise steadily, significantly impacting human health. Early prediction of pre-diabetic risks has emerged as a crucial public health concern in recent years. Machine learning methods have proven effective in enhancing prediction accuracy. However, existing approaches may lack interpretability regarding underlying mechanisms. Therefore, we aim to employ an interpretable machine learning approach utilizing nationwide cross-sectional data to predict pre-diabetic risk and quantify the impact of potential risks.

methodsThe LASSO regression algorithm was used to conduct feature selection from 30 factors, ultimately identifying nine non-zero coefficient features associated with pre-diabetes, including age, TG, TC, BMI, Apolipoprotein B, TP, leukocyte count, HDL-C, and hypertension. Various machine learning algorithms, including Extreme Gradient Boosting (XGBoost), Random Forest (RF), Support Vector Machine (SVM), Naive Bayes (NB), Artificial Neural Networks (ANNs), Decision Trees (DT), and Logistic Regression (LR), were employed to compare predictive performance. Employing an interpretable machine learning approach, we aimed to enhance the accuracy of pre-diabetes risk prediction and quantify the impact and significance of potential risks on pre-diabetes.

resultsFrom the China Health and Nutrition Survey (CHNS) data, a cohort of 8,277 individuals was selected, exhibiting a disease prevalence of 7.13%. The XGBoost model demonstrated superior performance with an AUC value of 0.939, surpassing RF, SVM, DT, ANNs, Naive Bayes, and LR models. Additionally, Shapley Additive Explanation (SHAP) analysis indicated that age, BMI, TC, ApoB, TG, hypertension, TP, HDL-C, and WBC may serve as risk factors for pre-diabetes.

conclusionThe constructed model comprises nine easily accessible predictive factors, which prove highly effective in forecasting the risk of pre-diabetes. Concurrently, we have quantified the specific impact of each predictive factor on the risk and ranked them based on their influence. This result may serve as a convenient tool for early identification of individuals at high risk of pre-diabetes, providing effective guidance for preventing the progression of pre-diabetes to T2DM.

Indexed as

Machine LearningPrediabetic StateAdultAgedAlgorithmsChinaCross-Sectional StudiesDiabetes Mellitus, Type 2FemaleHumansMaleMiddle AgedNutrition SurveysRisk AssessmentRisk FactorsLASSO regressionPre-diabetesPredictionShapley additive explanationXGBoost model

Identifiers

PMID40140819
PMCPMC11938594

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.