Evidence map›Paper›PMID 42679232›Full record

ArticleJMIR mental health2026

Large Language Model-Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation.

Florian Onur Kuhlmeier, Leon Hanschmann, Melina Rabe, Stefan Lüttke, Eva-Lotta Brakemeier, Alexander Maedche

Abstract read
In one paragraph

Article in JMIR mental health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Florian Onur KuhlmeierInstitute for Information Systems (WIN), Karlsruhe Institute of Technology, Kaiserstraße 89-93, Karlsruhe, Baden-Wurttemberg, 76133, Germany, 49 721 608-48380.ORCID http://orcid.org/0000-0002-1032-6982
Leon HanschmannInstitute for Information Systems (WIN), Karlsruhe Institute of Technology, Kaiserstraße 89-93, Karlsruhe, Baden-Wurttemberg, 76133, Germany, 49 721 608-48380.ORCID http://orcid.org/0000-0002-2002-0725
Melina RabeDepartment of Clinical Psychology and Psychotherapy, Universität Greifswald, Greifswald, Mecklenburg-Vorpommern, Germany.ORCID http://orcid.org/0009-0009-5602-2974
Stefan LüttkeChair of Clinical Child and Adolescent Psychology and Psychotherapy, Department of Psychology, Saarland University, Saarbrücken, Germany.ORCID http://orcid.org/0000-0002-9194-276X
Eva-Lotta BrakemeierDepartment of Clinical Psychology and Psychotherapy, Universität Greifswald, Greifswald, Mecklenburg-Vorpommern, Germany.ORCID http://orcid.org/0000-0001-9589-3697
Alexander MaedcheInstitute for Information Systems (WIN), Karlsruhe Institute of Technology, Kaiserstraße 89-93, Karlsruhe, Baden-Wurttemberg, 76133, Germany, 49 721 608-48380.ORCID http://orcid.org/0000-0001-6546-4816

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Mental health chatbots are increasingly used to support people with depressive symptoms, and large language models make these systems more flexible than rule-based chatbots. However, it remains unclear how well large language model-based chatbots deliver structured psychological interventions. Objective: This study examined how well a GPT-4o-based chatbot delivered a behavioral activation intervention for young people with depression using sessions with artificial users and clinical expert assessment. It also identified limitations and potential refinements. Methods: We implemented a GPT-4o (gpt-4o-2024-08-06; OpenAI)-based chatbot using a structured system prompt to deliver a single-session behavioral activation intervention for people with depression aged 14 to 29 years. We generated 48 sessions with GPT-4o-based artificial users derived from clinical vignettes varying across 7 characteristics. Ten clinical experts, either licensed psychotherapists or advanced psychotherapy trainees, independently assessed the sessions using the 14-item Quality of Behavioral Activation Scale (Q-BAS), rated from 0 to 6, supplemented by rating therapeutic capabilities, artificial user authenticity and difficulty, and qualitative feedback. Results: The chatbot completed all 7 intervention phases in every session. The mean holistic session quality rating was 3.94 (SD 1.23), and the mean Q-BAS rating was 4.03 (SD 1.18). Thirteen of 14 Q-BAS components exceeded the satisfactory threshold of 3. Ratings were highest for mood assessment (mean 5.42, SD 1.09) and activity planning (mean 4.98, SD 1.41) and lowest for explaining positive reinforcement (mean 2.92, SD 2.30) and supporting activity-mood monitoring (mean 3.02, SD 2.04). Therapeutic capability ratings were highest for message safety (mean 5.90, SD 0.37), message clarity (mean 5.56, SD 0.77), and objective, nonjudgmental communication (mean 5.17, SD 1.04) and lowest for therapeutic rapport (mean 4.12, SD 1.45) and natural conversation flow (mean 4.25, SD 1.42). Artificial users were rated below the scale midpoint for authenticity (mean 2.75, SD 1.41) and difficulty (mean 1.23, SD 1.46). Clinical experts described the chatbot as structured, clear, and safe but identified insufficient clinical reasoning as the main limitation, particularly in evaluating the therapeutic suitability and feasibility of activities, barriers, solution strategies, and rewards. Artificial users were often highly compliant, especially when identifying positive activities. Conclusions: In expert-rated sessions with artificial users, the chatbot delivered the behavioral activation intervention as intended and performed strongest on procedural components. It performed less well on positive reinforcement and activity-mood monitoring, indicating refinement needs in clinical reasoning, follow-up questioning, and evaluating whether proposed activities, plans, barriers, solution strategies, and rewards are therapeutically appropriate and feasible. The findings identify targets for improvement before testing with human users, while the artificial user design and expert ratings limit conclusions about real therapeutic interactions.

Indexed as

Behavior TherapyDepressionAdolescentAdultFemaleHumansLarge Language ModelsMaleMental Health TeletherapyYoung Adultartificial usersbehavioral activationclinical fidelitydepressiondigital mental health interventionslarge language modelsmental health chatbotsprompt engineeringyoung people

Identifiers

PMID42679232
PMCPMC13533316

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.