ArticleArchives of medical science : AMS2026
Machine learning predicts diabetes risk in high-risk populations: analysis of National Health and Nutrition Examination Survey data.
Article in Archives of medical science : AMS, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
5 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Introduction: This project intended to develop and validate a diabetes prediction model for high-risk populations based on machine learning algorithms. Material and methods: A total of 2,355 samples from the National Health and Nutrition Examination Survey (NHANES) database covering three cycles from 2013 to 2018 were included. The data were divided into training and testing sets in a 7 : 3 ratio. Nineteen risk prediction factors were selected as feature variables, including demographic baseline data, measurement data, medical history, and psychological health. Five machine learning models - decision tree, random forest (RF), multilayer perceptron (MLP), Adaboost, and Extreme Gradient Boosting (XGBoost) - were developed based on the data and variables mentioned above. Model performance was evaluated using accuracy, sensitivity, specificity, the area under curve (AUC) values of receiver operating characteristic (ROC) curves, and Matthews Correlation Coefficient (MCC) scores. Finally, the Shapley feature importance measurement tool was employed to select features in the optimal model. Results: The present work ultimately included 2,355 individuals at high risk of diabetes for analysis, with 260 cases of diabetes and 2,095 cases without diabetes. Among the five machine learning models established in this project., the RF and XGBoost models exhibited better overall performance compared to other models. In the test set, the RF model had an AUC of 0.896, accuracy of 0.784, sensitivity of 0.739, specificity of 0.849, and MCC of 0.418. The XGBoost model had corresponding values of AUC as 0.903, accuracy of 0.815, sensitivity of 0.962, and MCC of 0.443. According to the importance analysis of features in these two optimal models, waist circumference, age, BMI, gender, systolic blood pressure (SBP), diastolic blood pressure (DBP), education level, poverty income ratio (PIR), Patient Health Questionnaire (PHQ)-9 score, and race were the top ten key risk factors for diabetes in the high-risk population. Conclusions: The RF and XGBoost machine learning models demonstrated strong performance in predicting the occurrence of diabetes in high-risk populations. These models can aid in developing more precise intervention measures and personalized treatment plans to effectively reduce the incidence of diabetes and related risks in this population.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.