Evidence map›Paper›PMID 32442157›Full record

ArticleJournal of medical Internet research2020

Technical Metrics Used to Evaluate Health Care Chatbots: Scoping Review.

Alaa Abd-Alrazaq, Zeineb Safi, Mohannad Alajlani, Jim Warren, Mowafa Househ, Kerstin Denecke

Abstract readScoping Review
In one paragraph

Article in Journal of medical Internet research, 2020. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 52 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
52citing papers in PubMed, 2 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

52 citing papers in PubMed, 2 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Trial
  4. Trial
  5. Article
  6. Review
  7. Article
  8. Article
  9. Article
  10. Article
  11. Review
  12. Article
  13. Article
  14. Article
  15. Article
  16. Article
  17. Article
  18. Article
  19. Article
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Alaa Abd-AlrazaqCollege of Science and Engineering, Hamad Bin Khalifa University, Doha, Qatar.ORCID 0000-0001-7695-4626
Zeineb SafiCollege of Science and Engineering, Hamad Bin Khalifa University, Doha, Qatar.ORCID 0000-0003-0526-8949
Mohannad AlajlaniInstitute of Digital Healthcare, University of Warwick, Coventry, United Kingdom.ORCID 0000-0002-5691-7120
Jim WarrenSchool of Computer Science, University of Auckland, Auckland, New Zealand.ORCID 0000-0002-8660-8951
Mowafa HousehCollege of Science and Engineering, Hamad Bin Khalifa University, Doha, Qatar.ORCID 0000-0002-3648-6271
Kerstin DeneckeInstitute for Medical Informatics, Bern University of Applied Sciences, Bern, Switzerland.ORCID 0000-0001-6691-396X

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundDialog agents (chatbots) have a long history of application in health care, where they have been used for tasks such as supporting patient self-management and providing counseling. Their use is expected to grow with increasing demands on health systems and improving artificial intelligence (AI) capability. Approaches to the evaluation of health care chatbots, however, appear to be diverse and haphazard, resulting in a potential barrier to the advancement of the field.

objectiveThis study aims to identify the technical (nonclinical) metrics used by previous studies to evaluate health care chatbots.

methodsStudies were identified by searching 7 bibliographic databases (eg, MEDLINE and PsycINFO) in addition to conducting backward and forward reference list checking of the included studies and relevant reviews. The studies were independently selected by two reviewers who then extracted data from the included studies. Extracted data were synthesized narratively by grouping the identified metrics into categories based on the aspect of chatbots that the metrics evaluated.

resultsOf the 1498 citations retrieved, 65 studies were included in this review. Chatbots were evaluated using 27 technical metrics, which were related to chatbots as a whole (eg, usability, classifier performance, speed), response generation (eg, comprehensibility, realism, repetitiveness), response understanding (eg, chatbot understanding as assessed by users, word error rate, concept error rate), and esthetics (eg, appearance of the virtual agent, background color, and content).

conclusionsThe technical metrics of health chatbot studies were diverse, with survey designs and global usability metrics dominating. The lack of standardization and paucity of objective measures make it difficult to compare the performance of health chatbots and could inhibit advancement of the field. We suggest that researchers more frequently include metrics computed from conversation logs. In addition, we recommend the development of a framework of technical metrics with recommendations for specific circumstances for their inclusion in chatbot studies.

Indexed as

Artificial IntelligenceCommunicationDelivery of Health CareHumanschatbotsconversational agentsevaluationhealth caremetrics

Identifiers

PMID32442157
PMCPMC7305563

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.