ArticleActa informatica medica : AIM : journal of the Society for Medical Informatics of Bosnia & Herzegovina : casopis Drustva za medicinsku informatiku BiH2025
Machine Learning-Based Classification of Diabetes: Model Accuracy, Feature Importance, and Clinical Implications.
Article in Acta informatica medica : AIM : journal of the Society for Medical Informatics of Bosnia & Herzegovina : casopis Drustva za medicinsku informatiku BiH, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
5 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Diabetes mellitus (DM) is highly prevalent and often remains undiagnosed until complications appear, especially in low- and middle-income countries. Simple tools that use routinely collected clinical and demographic variables may support earlier identification of individuals at increased risk. Objective: This study aimed to build a supervised achine-learning model to classify individuals as diabetic or non-diabetic using a large publicly available dataset, and to identify which variables contributed most to the model decisions. Methods: We analysed a cleaned subset of 89,540 records from a Kaggle diabetes dataset. A multilayer perceptron artificial neural network (ANN) was trained and tested on separate subsets. Model performance was evaluated by overall accuracy and misclassification rates, and post-hoc variable importance scores were used to summarise the contribution of each predictor. Results: The ANN achieved an overall prediction accuracy of 96.8% in both the training and testing samples. Most records were correctly classified, although the error pattern suggested that non-diabetic cases were recognised more easily than diabetic cases. Blood glucose, HbA1c and body mass index (BMI) showed the highest importance values, whereas demographic and lifestyle variables contributed less to the classification. Conclusion: In this dataset, an ANN based on simple clinical and demographic variables was able to distinguish between diabetic and non-diabetic records with high internal accuracy and a plausible pattern of variable importance. The model could form the basis for a practical screening aid, but it requires external validation and further work on handling class imbalance and explainability before use in routine care.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.