Evidence map›Paper›PMID 42837528›Full record

SynthesisJMIR mental health2026

Effectiveness of Chatbots in Mental Health Screening and Assessment: Systematic Review.

Isabel Morales Gil, Inés Marti-Estevez, Sandra Doval, José Miguel Gutiérrez-Carrillo, Rocío Fausor, Silvia Arribas-García, Miguel Ángel Álvarez-Mon, Javier Quintero

Abstract readSystematic Review
In one paragraph

Synthesis in JMIR mental health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Isabel Morales GilUniversidad Internacional De La Rioja, Avenida de la Paz, Logroño, La Rioja, Spain, +34 941 20 97 43.ORCID http://orcid.org/0009-0007-5308-7698
Inés Marti-EstevezHospital Infantil Universitario Niño Jesús, Madrid, Spain.ORCID http://orcid.org/0009-0008-9749-8349
Sandra DovalUniversidad Internacional De La Rioja, Avenida de la Paz, Logroño, La Rioja, Spain, +34 941 20 97 43.ORCID http://orcid.org/0000-0002-1681-9016
José Miguel Gutiérrez-CarrilloGijón Hospital, Gijón, Spain.ORCID http://orcid.org/0009-0006-4177-7766
Rocío FausorValencian International University, Valencia, Spain.ORCID http://orcid.org/0000-0002-0309-3391
Silvia Arribas-GarcíaUniversidad Internacional De La Rioja, Avenida de la Paz, Logroño, La Rioja, Spain, +34 941 20 97 43.ORCID http://orcid.org/0000-0002-2747-5271
Miguel Ángel Álvarez-MonDepartment of Legal Medicine and Psychiatry, Complutense University, Madrid, Spain.ORCID http://orcid.org/0000-0002-1987-0394
Javier QuinteroDepartment of Legal Medicine and Psychiatry, Complutense University, Madrid, Spain.ORCID http://orcid.org/0000-0002-2491-8647

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Mental health disorders affect 970 million people globally; yet, over 50% do not access timely evaluation due to structural barriers and professional shortages. Chatbots and AI-based conversational agents have emerged as promising tools for mental health screening and assessment. Objective: This study systematically evaluated the effectiveness, accuracy, reliability, and acceptability of chatbots and AI-based conversational agents for mental health screening and assessment in adults. Methods: Systematic search conducted in May 2025 across PubMed/MEDLINE, PsycINFO, Scopus, and Web of Science (2019-2025), following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines. Eligible studies evaluated chatbots or AI for mental health screening/assessment in adults (≥18 y). Risk of bias was assessed using appropriate tools (Risk of Bias 2, Quality Assessment of Diagnostic Accuracy Studies-2, Joanna Briggs Institute checklists, and Mixed Methods Appraisal Tool). This systematic review was registered with PROSPERO (International Prospective Register of Systematic Reviews; CRD420251072392). Results: Eighteen studies (2021-2025) were included, with samples ranging from 20 to 3902 participants. Rule-based chatbots demonstrated high reliability (Cronbach α >0.85) and good acceptability (Acceptability of Intervention Measure >19/25). Generative models (large language models) achieved sensitivities of 0.84 to 0.93 and specificities of 0.80 to 0.96 for depression and anxiety, with correlations up to Conclusions: Chatbots and AI conversational agents demonstrate clinically relevant performance in mental health screening and assessment. However, safe implementation requires clear clinical protocols, professional supervision, integration with electronic health records, and active mitigation of algorithmic bias. These technologies should complement rather than replace clinical judgment.

Indexed as

Generative Artificial IntelligenceMass ScreeningMental DisordersHumansReproducibility of ResultsAIchatbotslarge language modelsmental healthpsychological assessmentscreening

Identifiers

PMID42837528
PMCPMC13641319

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.