Evidence map›Paper›PMID 41592221›Full record

ReviewJMIR AI2026

Large Language Model-Based Chatbots and Agentic AI for Mental Health Counseling: Systematic Review of Methodologies, Evaluation Frameworks, and Ethical Safeguards.

Ha Na Cho, Jiayuan Wang, Di Hu, Kai Zheng

Abstract readReview
In one paragraph

Review in JMIR AI, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Trial
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Ha Na ChoDepartment of Informatics, University of California, Irvine, Irvine, CA, United States.ORCID https://orcid.org/0000-0001-8033-6644
Jiayuan WangDepartment of Informatics, University of California, Irvine, Irvine, CA, United States.ORCID https://orcid.org/0009-0000-2144-4610
Di HuDepartment of Informatics, University of California, Irvine, Irvine, CA, United States.ORCID https://orcid.org/0000-0002-2842-1478
Kai ZhengDepartment of Informatics, University of California, Irvine, Irvine, CA, United States.ORCID https://orcid.org/0000-0003-4121-4948

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundLarge language model (LLM)-based chatbots have rapidly emerged as tools for digital mental health (MH) counseling. However, evidence on their methodological quality, evaluation rigor, and ethical safeguards remains fragmented, limiting interpretation of clinical readiness and deployment safety.

objectiveThis systematic review aimed to synthesize the methodologies, evaluation practices, and ethical or governance frameworks of LLM-based chatbots developed for MH counseling and to identify gaps affecting validity, reproducibility, and translation.

methodsWe searched Google Scholar, PubMed, IEEE Xplore, and ACM Digital Library for studies published between January 2020 and May 2025. Eligible studies reported original development or empirical evaluation of LLM-driven MH counseling chatbots. We excluded studies that did not involve LLM-based conversational agents, were not focused on counseling or supportive MH communication, or lacked evaluable system outputs or outcomes. Screening and data extraction were conducted in Covidence (Veritas Health Innovation) following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines. Study quality was appraised using a structured traffic light framework across 5 methodological domains (design, dataset reporting, evaluation metrics, external validation, and ethics), with an overall judgment derived across domains. We used narrative synthesis with descriptive aggregation to summarize methodological trends, evaluation metrics, and governance considerations.

resultsTwenty studies met the inclusion criteria. GPT-based models (GPT-2/3/4) were used in 45% (9/20) of studies, while 90% (18/20) used fine-tuned or domain-adaptation models such as LLaMa, ChatGLM, or Qwen. Reported deployment types were not mutually exclusive; standalone apps were most common (18/20, 90%), and some systems were also implemented as virtual agents (4/20, 20%) or delivered via existing platforms (2/20, 10%). Evaluation approaches were frequently mixed, with qualitative assessment (13/20, 65%), such as thematic analysis or rubric-based scoring, often complemented by quantitative language metrics (18/20, 90%), including BLEU (Bilingual Evaluation Understudy), ROUGE (Recall-Oriented Understudy for Gisting Evaluation), or perplexity. Quality appraisal indicated consistently low risk for dataset reporting and evaluation metrics, but recurring limitations were observed in external validation and reporting on ethics and safety, including incomplete documentation of safety safeguards and governance practices. No included study reported registered randomized controlled trials or independent clinical validation in real-world care settings.

conclusionsLLM-based MH counseling chatbots show promise for scalable and personalized support, but current evidence is limited by heterogeneous study designs, minimal external validation, and inconsistent reporting of safety and governance practices. Future work should prioritize clinically grounded evaluation frameworks, transparent reporting of model and prompt configurations, and stronger validation using standardized outcomes to support safe, reliable, and regulatory-ready deployment.

Indexed as

conversational agentdigital healthdigital mental health interventionlarge language model chatbotspersonalized health care

Identifiers

PMID41592221
PMCPMC13032092

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.