Evidence map›Paper›PMID 41853071›Full record

SynthesisFrontiers in psychiatry2026

Artificial intelligence in mental health care: a scoping review of reviews.

Mohammad S Abu-Mahfouz, Sarah AlFehaid, Hala M Burqan, Rabie Adel El Arab

Abstract readSystematic Review
In one paragraph

Synthesis in Frontiers in psychiatry, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 7 papers.

0numbers the graph read from it
0cells of the map it votes in
7citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

7 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Mohammad S Abu-MahfouzAlmoosa College of Health Sciences, Al Ahsa, Saudi Arabia.
Sarah AlFehaidAlmoosa College of Health Sciences, Al Ahsa, Saudi Arabia.
Hala M BurqanAlmoosa College of Health Sciences, Al Ahsa, Saudi Arabia.
Rabie Adel El ArabAlmoosa College of Health Sciences, Al Ahsa, Saudi Arabia.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Artificial intelligence (AI) is rapidly entering mental health care, but most models remain proof-of-concept, with limited external validation and substantial risk of overfitting. Methods: This scoping review of reviews adhered to the PRISMA-ScR checklist and Joanna Briggs Institute guidance. We searched MEDLINE, Embase, PsycINFO, and IEEE Xplore. Eligible publications encompassed systematic, scoping, narrative, integrative, meta-analytic, and patent reviews. Findings were synthesised thematically. Results: Thirty-one reviews were included. Evidence concentrated on depression and anxiety; schizophrenia, bipolar disorder, perinatal mental health, autism spectrum conditions, older adults, nurses, and allied professionals were under-represented. Across screening, diagnosis/classification, and risk prediction, high accuracy was frequently reported under internal validation; in prior syntheses, typical internal AUCs clustered around ≈0.80-0.88 whereas externally or prospectively validated performance was scarce and typically attenuated. Signals were strongest for narrow, feedback-rich tasks, with greater decay for general-purpose models and longer prediction horizons. Conversational agents produced small-to-moderate short-term improvements in depressive symptoms (SMD ≈0.2-0.6); effects for anxiety and stress were smaller or inconsistent and varied with comparator stringency, follow-up (≤8-12 weeks vs longer), and the degree of human guidance. Most chatbot evaluations were short and small-scale, with few randomized or pragmatic trials and limited data on durability beyond 12 weeks. Real-world implementation was limited; several reviews identified usability and electronic health-record integration as prerequisites for adoption, and explainability alone rarely conferred actionability without clinician training. Ethical readiness was incomplete: privacy and bias were commonly discussed, but accountability, post-deployment monitoring, and crisis-escalation protocols were inconsistently specified. Economic evaluations were uncommon and rarely accounted for integration, maintenance, or re-training costs. Workforce outcomes (literacy, confidence, readiness) were infrequently measured. Internal and external metrics were not pooled. Conclusions: AI applications span the mental-health care continuum but remain early in translation. Performance that appears strong under internal validation often attenuates on external or prospective testing; symptomatic gains are concentrated in depression/anxiety and may diminish over longer follow-up; and adoption is constrained by usability, EHR integration, and incomplete governance. The cross-review signal highlights consistent gaps in accountability, post-deployment monitoring and crisis escalation, equity reporting, workforce readiness, and life-cycle economics (including integration, monitoring, and re-training). Addressing these gaps through externally validated and monitored deployments, routine content/guardrail audits for chatbots with human escalation, predefined subgroup performance and bias auditing, and implementation strategies that pair explainability with clinician training and measure workforce endpoints would better align the evidence base with safe, effective, and sustainable clinical use.

Indexed as

anxiety disordersartificial intelligencediagnosisdigital therapeuticsmental healthmood disorderspredictive modellingscreening

Identifiers

PMID41853071
PMCPMC12993279

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.