Evidence map›Paper›PMID 41775771›Full record

ArticleScientific reports2026

A hybrid SMOTE and Gaussian mixture model based optimized XGBoost framework for bipolar disorder detection.

Santosh Kumar, Deeksha Kumari, Arvind Panwar, Shrddha Sagar, Lukas Herout, Hamidreza Namazi, Nitesh Singh Bhati

Abstract read
In one paragraph

Article in Scientific reports, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Santosh KumarSchool of Computing Science and Engineering, Galgotias University, Greater Noida, Uttar Pradesh, India.
Deeksha KumariSchool of Computing Science and Engineering, Galgotias University, Greater Noida, Uttar Pradesh, India.
Arvind PanwarSchool of Computing Science and Engineering, Galgotias University, Greater Noida, Uttar Pradesh, India. arvind.nice3@gmail.com.
Shrddha SagarSchool of Computing Science and Engineering, Galgotias University, Greater Noida, Uttar Pradesh, India.
Lukas HeroutDepartment of Informatics, Skoda Auto University, Mlada Boleslav, Czech Republic.
Hamidreza NamaziSchool of Engineering, Monash University, Selangor, Malaysia. hamidreza.namazi@monash.edu.
Nitesh Singh BhatiDepartment of Computer Science and Engineering, School of ICT, Gautam Buddha University Greater Noida, Greater Noida, Uttar Pradesh, India.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

The identification of bipolar disorder (BD), a severe psychiatric condition characterized by recurrent mood fluctuations, remains challenging due to substantial inter-individual variability, symptom overlap with other mental disorders, and imbalanced clinical data. Delayed or inaccurate diagnosis often leads to inappropriate treatment strategies and adverse clinical outcomes, highlighting the need for reliable, data-driven decision-support tools. In this study, we propose a robust hybrid machine learning framework that integrates class balancing, latent subgroup discovery, and ensemble learning to improve the accuracy and consistency of BD identification from tabular clinical data. The framework applies the Synthetic Minority Over-sampling Technique (SMOTE) exclusively to the training data to address class imbalance, followed by Gaussian Mixture Model (GMM) based clustering to uncover latent patient subgroups and generate informative probabilistic features. These enriched features are subsequently used to train an optimized Extreme Gradient Boosting (XGBoost) classifier. Experimental evaluation on an independent test set demonstrates that the proposed model achieves 93% accuracy, 97% sensitivity (recall), 93% precision, 95% F1-score, and 79% specificity. When evaluated under identical experimental conditions, the proposed framework consistently outperforms baseline classifiers, including Support Vector Machine, Decision Tree, Logistic Regression, and Random Forest, with performance improvements ranging from 6 to 12%, depending on the comparator. The results indicate that combining SMOTE-based data balancing, GMM-driven latent feature enrichment, and gradient-boosted decision trees yields a scalable, interpretable, and clinically relevant decision-support system. This study supports the adoption of hybrid, data-driven approaches for early BD screening and personalized treatment planning in psychiatric healthcare settings.

Indexed as

Bipolar DisorderBoosting Machine Learning AlgorithmsClassification AlgorithmsHumansMachine LearningNormal DistributionBipolar disorderGaussian mixture model (GMM)Mental health diagnosticsSMOTEXGBoost

Identifiers

PMID41775771
PMCPMC13066589

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.