Evidence map›Paper›PMID 39903463›Full record

SynthesisJAMA network open2025

Large Language Models for Chatbot Health Advice Studies: A Systematic Review.

Bright Huo, Amy Boyle, Nana Marfo, Wimonchat Tangamornsuksan, Jeremy P Steen, Tyler McKechnie, Yung Lee, Julio Mayol, Stavros A Antoniou, Arun James Thirunavukarasu and 3 more

Erratum issuedAbstract readSystematic Review
In one paragraph

Synthesis in JAMA network open, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. An erratum has been issued. Cited by 120 papers, 5 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
120citing papers in PubMed, 5 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

120 citing papers in PubMed, 5 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Pooled it
  4. Guideline
  5. Pooled it
  6. Trial
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Validating LLM judges for automated oversight of patient communication.medRxiv : the preprint server for health sciences · 2026
    Article
  13. Article
  14. Article
  15. Article
  16. Article
  17. Article
  18. Article
  19. Article
  20. Article

60 more citing papers are in PubMed but not listed here.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

13 authors.

Bright HuoDivision of General Surgery, Department of Surgery, McMaster University, Hamilton, Ontario, Canada.
Amy BoyleMichael G. DeGroote School of Medicine, McMaster University, Hamilton, Ontario, Canada.
Nana MarfoH. Ross University School of Medicine, Miramar, Florida.
Wimonchat TangamornsuksanDepartment of Health Research Methods, Evidence, and Impact, McMaster University, Hamilton, Ontario, Canada.
Jeremy P SteenInstitute of Health Policy, Management and Evaluation, University of Toronto, Toronto, Ontario, Canada.
Tyler McKechnieDivision of General Surgery, Department of Surgery, McMaster University, Hamilton, Ontario, Canada.
Yung LeeDivision of General Surgery, Department of Surgery, McMaster University, Hamilton, Ontario, Canada.
Julio MayolHospital Clinico San Carlos, IdISSC, Universidad Complutense de Madrid, Madrid, Spain.
Stavros A AntoniouDepartment of Surgery, Papageorgiou General Hospital, Thessaloniki, Greece.
Arun James ThirunavukarasuOxford University Clinical Academic Graduate School, University of Oxford, Oxford, United Kingdom.
Stephanie SangerHealth Science Library, Faculty of Health Sciences, McMaster University, Hamilton, Ontario, Canada.
Karim RamjiDivision of General Surgery, Department of Surgery, McMaster University, Hamilton, Ontario, Canada.
Gordon GuyattDepartment of Clinical Epidemiology and Biostatistics, McMaster University, Hamilton, Ontario, Canada.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Importance: There is much interest in the clinical integration of large language models (LLMs) in health care. Many studies have assessed the ability of LLMs to provide health advice, but the quality of their reporting is uncertain. Objective: To perform a systematic review to examine the reporting variability among peer-reviewed studies evaluating the performance of generative artificial intelligence (AI)-driven chatbots for summarizing evidence and providing health advice to inform the development of the Chatbot Assessment Reporting Tool (CHART). Evidence Review: A search of MEDLINE via Ovid, Embase via Elsevier, and Web of Science from inception to October 27, 2023, was conducted with the help of a health sciences librarian to yield 7752 articles. Two reviewers screened articles by title and abstract followed by full-text review to identify primary studies evaluating the clinical accuracy of generative AI-driven chatbots in providing health advice (chatbot health advice studies). Two reviewers then performed data extraction for 137 eligible studies. Findings: A total of 137 studies were included. Studies examined topics in surgery (55 [40.1%]), medicine (51 [37.2%]), and primary care (13 [9.5%]). Many studies focused on treatment (91 [66.4%]), diagnosis (60 [43.8%]), or disease prevention (29 [21.2%]). Most studies (136 [99.3%]) evaluated inaccessible, closed-source LLMs and did not provide enough information to identify the version of the LLM under evaluation. All studies lacked a sufficient description of LLM characteristics, including temperature, token length, fine-tuning availability, layers, and other details. Most studies (136 [99.3%]) did not describe a prompt engineering phase in their study. The date of LLM querying was reported in 54 (39.4%) studies. Most studies (89 [65.0%]) used subjective means to define the successful performance of the chatbot, while less than one-third addressed the ethical, regulatory, and patient safety implications of the clinical integration of LLMs. Conclusions and Relevance: In this systematic review of 137 chatbot health advice studies, the reporting quality was heterogeneous and may inform the development of the CHART reporting standards. Ethical, regulatory, and patient safety considerations are crucial as interest grows in the clinical integration of LLMs.

Indexed as

Artificial IntelligenceGenerative Artificial IntelligenceHumansLarge Language Models

Identifiers

PMID39903463
PMCPMC11795331

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.