SynthesisJAMA network open2025
Large Language Models for Chatbot Health Advice Studies: A Systematic Review.
Synthesis in JAMA network open, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. An erratum has been issued. Cited by 120 papers, 5 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
120 citing papers in PubMed, 5 syntheses or guidelines pooled it.
- Application of Large Language Models in Chronic Disease Care: Mixed Methods Systematic Review and Thematic Synthesis.Journal of medical Internet research · 2026Pooled it
- Impact of Large Language Model-Based AI Tools on Physician-Patient Communication: Systematic Review and Meta-Analysis.Journal of medical Internet research · 2026Pooled it
- Large language models for primary care ophthalmic education: a systematic review.Frontiers in medicine · 2026Pooled it
- Reporting guideline for Chatbot Health Advice studies: the CHART statement.BMC medicine · 2025Guideline
- Evaluating the Performance of ChatGPT on Board-Style Examination Questions in Ophthalmology: A Meta-Analysis.Journal of medical systems · 2025Pooled it
- Preliminary Evaluation of a Large Language Model-Powered Chatbot for Osteoporosis Self-Management Education: Formative Randomized Controlled Trial.JMIR formative research · 2026Trial
- Telephone triage in urgent unscheduled primary care in 16 European countries: a cross-national questionnaire-based expert study.Scandinavian journal of primary health care · 2026Article
- Real-world use and evaluation of a generative AI chatbot for Parkinson's disease information: a prospective observational study.The Lancet regional health. Europe · 2026Article
- ChatGPT Health and ChatGPT Plus in Urogynecology: A Blinded Comparison of Response Quality, Guideline Concordance, and Safety.International urogynecology journal · 2026Article
- Audiologist-Guided Multimodal AI for Pure-Tone Audiometry and Tympanometry Interpretation and Reporting.Journal of medical systems · 2026Article
- Conversational AI in Hereditary Cancer Care: Sociotechnical Study of Responsible Design Requirements.JMIR human factors · 2026Article
- Validating LLM judges for automated oversight of patient communication.medRxiv : the preprint server for health sciences · 2026Article
- Consensus framework for the validation of generative AI: call for collaborators on the Validation Accords.Nature medicine · 2026Article
- Evaluating the quality of artificial intelligence responses to psoriasis-related clinical and patient questions: a comparative study of ChatGPT, Gemini, and Microsoft Copilot.Proceedings (Baylor University. Medical Center) · 2026Article
- Can Artificial Intelligence Substitute Dental Public Health Expertise? A Competency-Based Analysis Across Core DPH Domains.Journal of public health dentistry · 2026Article
- Comparative performance of artificial intelligence chatbots in patient education for robot-assisted radical prostatectomy: quality, transparency and readability.Journal of robotic surgery · 2026Article
- Critical Care-Specific vs General-Purpose Large Language Models in Emergency Intensive Care Unit Diagnosis: Single-Center Retrospective Paired Comparative Study.Journal of medical Internet research · 2026Article
- Comparative Performance of AI Models and Clinicians in Evidence-Based Cardiovascular Disease Management for People Living With HIV: Comparative Study.Journal of medical Internet research · 2026Article
- Reliability and readability of AI chatbot responses to patient questions about robot-assisted radical cystectomy.Journal of robotic surgery · 2026Article
- Multiturn Large Language Model-Based Conversational Agents for Patients With Cancer and Caregivers: Scoping Review.JMIR cancer · 2026Article
60 more citing papers are in PubMed but not listed here.
Corrections and comments
- Erratum issued
Authors and funding
13 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Importance: There is much interest in the clinical integration of large language models (LLMs) in health care. Many studies have assessed the ability of LLMs to provide health advice, but the quality of their reporting is uncertain. Objective: To perform a systematic review to examine the reporting variability among peer-reviewed studies evaluating the performance of generative artificial intelligence (AI)-driven chatbots for summarizing evidence and providing health advice to inform the development of the Chatbot Assessment Reporting Tool (CHART). Evidence Review: A search of MEDLINE via Ovid, Embase via Elsevier, and Web of Science from inception to October 27, 2023, was conducted with the help of a health sciences librarian to yield 7752 articles. Two reviewers screened articles by title and abstract followed by full-text review to identify primary studies evaluating the clinical accuracy of generative AI-driven chatbots in providing health advice (chatbot health advice studies). Two reviewers then performed data extraction for 137 eligible studies. Findings: A total of 137 studies were included. Studies examined topics in surgery (55 [40.1%]), medicine (51 [37.2%]), and primary care (13 [9.5%]). Many studies focused on treatment (91 [66.4%]), diagnosis (60 [43.8%]), or disease prevention (29 [21.2%]). Most studies (136 [99.3%]) evaluated inaccessible, closed-source LLMs and did not provide enough information to identify the version of the LLM under evaluation. All studies lacked a sufficient description of LLM characteristics, including temperature, token length, fine-tuning availability, layers, and other details. Most studies (136 [99.3%]) did not describe a prompt engineering phase in their study. The date of LLM querying was reported in 54 (39.4%) studies. Most studies (89 [65.0%]) used subjective means to define the successful performance of the chatbot, while less than one-third addressed the ethical, regulatory, and patient safety implications of the clinical integration of LLMs. Conclusions and Relevance: In this systematic review of 137 chatbot health advice studies, the reporting quality was heterogeneous and may inform the development of the CHART reporting standards. Ethical, regulatory, and patient safety considerations are crucial as interest grows in the clinical integration of LLMs.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.