ArticleJournal of medical Internet research2020
Technical Metrics Used to Evaluate Health Care Chatbots: Scoping Review.
Article in Journal of medical Internet research, 2020. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 52 papers, 2 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
52 citing papers in PubMed, 2 syntheses or guidelines pooled it.
- A systematic review of artificial intelligence chatbots for promoting physical activity, healthy diet, and weight loss.The international journal of behavioral nutrition and physical activity · 2021Pooled it
- Voice-Based Conversational Agents for the Prevention and Management of Chronic and Mental Health Conditions: Systematic Literature Review.Journal of medical Internet research · 2021Pooled it
- The Impact of Chatbot Type and Normative Messaging on Chatbot Usage Intention Based on the Health Technology Acceptance Model: Randomized Controlled Trial.Journal of medical Internet research · 2026Trial
- Mental Health Chatbot for Young Adults With Depressive Symptoms During the COVID-19 Pandemic: Single-Blind, Three-Arm Randomized Controlled Trial.Journal of medical Internet research · 2022Trial
- User needs and design opportunities for a conversational agent for tuberculosis treatment: A mixed-methods study.PLOS digital health · 2026Article
- Metrics Used for the Evaluation of Chatbots Providing Cancer Genetic Risk Assessment and Education: Systematic Review.JMIR AI · 2026Review
- Assessment of a Digital Health Platform Using Web Analytics and User Experience Measurements: Quantitative Study Based on RE-AIM.Journal of medical Internet research · 2026Article
- Design, Development, and Evaluation of Multimodal Conversational Agents for Health Data Registration and Monitoring: Framework Proposal and Pilot Exploratory Study.Healthcare (Basel, Switzerland) · 2026Article
- Unlocking AI Chatbot Potential in Healthcare: Trust-Enhanced DeLone & McLean IS Success Model.Healthcare (Basel, Switzerland) · 2026Article
- Large language models for psychosocial risk assessment: A multi-method evaluation across suicide, intimate partner violence, and substance misuse.PLOS digital health · 2026Article
- From Tool to Agent: A Semi-Systematic Review of Human-AI Alignment and a Proposed Tiered Healing Ecosystem for Mental Health.Healthcare (Basel, Switzerland) · 2026Review
- Developing a Service Quality Index System for AI Health Care Chatbots: Mixed Methods Study.Journal of medical Internet research · 2026Article
- What can chatbot conversations reveal about vaccine concerns? An observational topic modelling study for public health infoveillance.BMJ digital health & AI · 2026Article
- My diabetes care: an AI-based mobile app with conversational agent for type 2 diabetes self-management.Scientific reports · 2025Article
- Paediatric rare diseases: Can large language models assist off-label prescribing?British journal of clinical pharmacology · 2025Article
- Digital Health Interventions for Depression and Anxiety in Low- and Middle-Income Countries: Rapid Scoping Review.JMIR mental health · 2025Article
- Parents' information needs and perceptions of chatbots regarding self medicating their children.Scientific reports · 2025Article
- LLM-Based Response Generation for Korean Adolescents: A Study Using the NAVER Knowledge iN Q&A Dataset with RAG.Healthcare informatics research · 2025Article
- Developing Effective Frameworks for Large Language Model-Based Medical Chatbots: Insights From Radiotherapy Education With ChatGPT.JMIR cancer · 2025Article
- Development of Chatbot-Based Oral Health Care for Young Children and Evaluation of its Effectiveness, Usability, and Acceptability: Mixed Methods Study.JMIR pediatrics and parenting · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundDialog agents (chatbots) have a long history of application in health care, where they have been used for tasks such as supporting patient self-management and providing counseling. Their use is expected to grow with increasing demands on health systems and improving artificial intelligence (AI) capability. Approaches to the evaluation of health care chatbots, however, appear to be diverse and haphazard, resulting in a potential barrier to the advancement of the field.
objectiveThis study aims to identify the technical (nonclinical) metrics used by previous studies to evaluate health care chatbots.
methodsStudies were identified by searching 7 bibliographic databases (eg, MEDLINE and PsycINFO) in addition to conducting backward and forward reference list checking of the included studies and relevant reviews. The studies were independently selected by two reviewers who then extracted data from the included studies. Extracted data were synthesized narratively by grouping the identified metrics into categories based on the aspect of chatbots that the metrics evaluated.
resultsOf the 1498 citations retrieved, 65 studies were included in this review. Chatbots were evaluated using 27 technical metrics, which were related to chatbots as a whole (eg, usability, classifier performance, speed), response generation (eg, comprehensibility, realism, repetitiveness), response understanding (eg, chatbot understanding as assessed by users, word error rate, concept error rate), and esthetics (eg, appearance of the virtual agent, background color, and content).
conclusionsThe technical metrics of health chatbot studies were diverse, with survey designs and global usability metrics dominating. The lack of standardization and paucity of objective measures make it difficult to compare the performance of health chatbots and could inhibit advancement of the field. We suggest that researchers more frequently include metrics computed from conversation logs. In addition, we recommend the development of a framework of technical metrics with recommendations for specific circumstances for their inclusion in chatbot studies.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.