Evidence map›Paper›PMID 41401240›Full record

SynthesisJournal of medical Internet research2025

Generative AI Mental Health Chatbots as Therapeutic Tools: Systematic Review and Meta-Analysis of Their Role in Reducing Mental Health Issues.

Qiyang Zhang, Renwen Zhang, Yiying Xiong, Yuan Sui, Chang Tong, Fu-Hung Lin

Abstract readSystematic ReviewMeta-Analysis
In one paragraph

Synthesis in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 18 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
18citing papers in PubMed, 2 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

18 citing papers in PubMed, 2 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Trial
  4. Article
  5. Article
  6. Article
  7. Article
  8. Artificial Intelligence in Spiritual Care: Modified Delphi Study.Journal of medical Internet research · 2026
    Article
  9. Article
  10. Article
  11. Article
  12. Review
  13. Article
  14. Article
  15. Article
  16. Review
  17. Review
  18. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Qiyang ZhangDepartment of Educational Advancement, Duke-NUS Medical School, 8 College Road, Singapore, 169857, Singapore, 65 66012186.ORCID http://orcid.org/0000-0001-7474-2435
Renwen ZhangWee Kim Wee School of Communication and Information, Nanyang Technological University, Singapore, Singapore.ORCID http://orcid.org/0000-0002-7636-9598
Yiying XiongSchool of Education, Johns Hopkins University, Baltimore, MD, United States.ORCID http://orcid.org/0000-0002-0887-5587
Yuan SuiSchool of Education, Johns Hopkins University, Baltimore, MD, United States.ORCID http://orcid.org/0009-0005-0917-1847
Chang TongSchool of Education, Johns Hopkins University, Baltimore, MD, United States.ORCID http://orcid.org/0009-0007-4599-5610
Fu-Hung LinSchool of Education, Johns Hopkins University, Baltimore, MD, United States.ORCID http://orcid.org/0009-0005-1766-2873

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: In recent years, artificial intelligence (AI) has driven the rapid development of AI mental health chatbots. Most current reviews investigated the effectiveness of rule-based or retrieval-based chatbots. To date, there is no comprehensive review that systematically synthesizes the effect of generative AI (GenAI) chatbot's impact on mental health. Objective: This review aims to (1) narratively synthesize existing GenAI mental health chatbots' technical features, treatment and research designs, and sample characteristics through a systematic review of quantitative studies and (2) quantify the effectiveness and key moderators of these rigorously designed trials on GenAI mental health chatbots through a meta-analysis of only randomized controlled trials (RCTs). Methods: The search strategy includes 11 database searching, backward citation tracking, and a manual ad hoc search to update literature. This thorough literature search, completed in March 2025, returned 5555 records for screening. The systematic review included studies that (1) used generative or hybrid (rule/retrieval-based and generative) AI-based chatbots to deliver interventions and (2) quantitatively measured mental health-related outcomes. The meta-analysis has additional inclusion criteria: (1) studies must be RCTs, (2) must measure negative mental health issues, (3) the comparison group must not have chatbot features, and (4) must provide enough statistics for effect size calculation. We followed the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) checklist and registered the protocol retrospectively during the revision process (September 18, 2025). In meta-regression, data were synthesized in R software using a random-effects model. Results: The narrative synthesis of 26 studies revealed that (1) GenAI chatbot interventions mostly took place in non-WEIRD countries (non-Western, Educated, Industrialized, Rich, and Democratic) and (2) there is a lack of studies focusing on young children and older adults. The meta-analysis of 14 RCTs showed a statistically significant effect (effect size [ES]=0.30, P=.047, N=6314, 95% CI 0.004, 0.59, 95% prediction interval [PI] -0.85, 1.67), which means that GenAI chatbots are, on average, effective in reducing negative mental health issues, such as depression, anxiety, among others. We found that social-oriented chatbots (ie, those that mainly provide social interactions) are more effective than task-oriented programs (ie, those that assist with specific tasks). Risk of bias in the nonrandomized studies and RCTs was assessed using Cochrane ROBINS-I (Risk Of Bias In Non-randomised Studies - of Interventions) and RoB2 (revised Cochrane risk-of-bias tool for randomized trials), respectively, indicating a moderate amount of risk. One main limitation of this meta-analysis is the small number of studies (n=14) included. Conclusions: By identifying research gaps, we suggest that future researchers investigate user groups such as adolescents and older adults, outcomes other than depression and anxiety, cultural adaptations in non-WEIRD countries, ways to streamline chatbots in usual care practices, and explore applications in diverse settings. More importantly, we cannot ignore GenAI chatbots' risks while acknowledging their promise. This review also emphasized several ethical implications.

Indexed as

Artificial IntelligenceMental DisordersMental HealthGenerative Artificial IntelligenceHumansRandomized Controlled Trials as TopicAI chatbotartificial intelligenceconversational agentsgenerative artificial intelligencemental healthmeta-analysissystematic review

Identifiers

PMID41401240
PMCPMC12707440

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.