SynthesisJAMA2025
Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review.
Synthesis in JAMA, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to trial NCT06997107 (AI-CARE), which is not on this map. Cited by 357 papers, 5 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
AI-CARE: AI to Create Accessible & Reliable Patient Education Materials
Who cites it
357 citing papers in PubMed, 5 syntheses or guidelines pooled it.
- Generative large language models in the clinical management of Alzheimer's disease and mild cognitive impairment.Neurological sciences : official journal of the Italian Neurological Society and of the Italian Society of Clinical Neurophysiology · 2026Pooled it
- Research on patient-facing chatbots based on large language models in the care of older people: a living systematic review.European geriatric medicine · 2026Pooled it
- Testing and evaluation of generative large language models in electronic health record applications: a systematic review.Journal of the American Medical Informatics Association : JAMIA · 2026Pooled it
- LLM-assisted systematic review of large language models in clinical medicine.Nature medicine · 2026Pooled it
- Large language models for simplifying radiology reports: a systematic review and meta-analysis of patient, public, and clinician evaluations.The Lancet. Digital health · 2026Pooled it
- Preliminary Evaluation of a Large Language Model-Powered Chatbot for Osteoporosis Self-Management Education: Formative Randomized Controlled Trial.JMIR formative research · 2026Trial
- Large language models enhance diagnostic reasoning of medical students in rheumatology: a randomized controlled trial.BMC medical education · 2026Trial
- An LLM chatbot to facilitate primary-to-specialist care transitions: a randomized controlled trial.Nature medicine · 2026Trial
- Misaligned by design: evaluating and deploying generative AI for the real-world conditions of primary care: a Nordic and European perspective.Scandinavian journal of primary health care · 2026Article
- Introduction to Concepts in Artificial Intelligence and Machine Learning for Pharmacoepidemiologists: Large Language Models.Pharmacoepidemiology and drug safety · 2026Article
- Evaluating the role of ChatGPT in patient questions regarding drug allergies.The World Allergy Organization journal · 2026Article
- Evaluating the Accuracy of Large Language Models in Risk-of-Bias Assessment Using Version 2 of the Cochrane Risk-of-Bias Tool for Randomized Trials: Exploratory Feasibility Study.Journal of medical Internet research · 2026Article
- Cloud-Based and Locally Deployed Language Models in Nursing and Health Care: An AI Act-Aligned Framework.JMIR medical informatics · 2026Article
- Validating LLM judges for automated oversight of patient communication.medRxiv : the preprint server for health sciences · 2026Article
- LungGPT: A unified multimodal system for interpretable diagnosis and clinical decision support of respiratory diseases.Cell reports. Medicine · 2026Article
- Large Language Models for Distress Rating in Korean Psycho-Oncology Interviews: Exploratory Clinician-Benchmarked Evaluation Study.Journal of medical Internet research · 2026Article
- Generative artificial intelligence to augment ethical problem solving in ophthalmology: GPT-5.1 versus a human ethicist.International ophthalmology · 2026Article
- Unmasking bias in the evidence ecosystem: a panoramic analysis of 311,751 meta-analyses using an artificial intelligence agent-based approach.Journal of global health · 2026Article
- Article
- Evaluating a Guideline-Integrated Clinical Interaction Framework Vs a Standard Large Language Model Interaction for Dietary Recommendations in Recurrent Urolithiasis: In Silico Study.Journal of medical Internet research · 2026Article
297 more citing papers are in PubMed but not listed here.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
19 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Importance: Large language models (LLMs) can assist in various health care activities, but current evaluation approaches may not adequately identify the most useful application areas. Objective: To summarize existing evaluations of LLMs in health care in terms of 5 components: (1) evaluation data type, (2) health care task, (3) natural language processing (NLP) and natural language understanding (NLU) tasks, (4) dimension of evaluation, and (5) medical specialty. Data Sources: A systematic search of PubMed and Web of Science was performed for studies published between January 1, 2022, and February 19, 2024. Study Selection: Studies evaluating 1 or more LLMs in health care. Data Extraction and Synthesis: Three independent reviewers categorized studies via keyword searches based on the data used, the health care tasks, the NLP and NLU tasks, the dimensions of evaluation, and the medical specialty. Results: Of 519 studies reviewed, published between January 1, 2022, and February 19, 2024, only 5% used real patient care data for LLM evaluation. The most common health care tasks were assessing medical knowledge such as answering medical licensing examination questions (44.5%) and making diagnoses (19.5%). Administrative tasks such as assigning billing codes (0.2%) and writing prescriptions (0.2%) were less studied. For NLP and NLU tasks, most studies focused on question answering (84.2%), while tasks such as summarization (8.9%) and conversational dialogue (3.3%) were infrequent. Almost all studies (95.4%) used accuracy as the primary dimension of evaluation; fairness, bias, and toxicity (15.8%), deployment considerations (4.6%), and calibration and uncertainty (1.2%) were infrequently measured. Finally, in terms of medical specialty area, most studies were in generic health care applications (25.6%), internal medicine (16.4%), surgery (11.4%), and ophthalmology (6.9%), with nuclear medicine (0.6%), physical medicine (0.4%), and medical genetics (0.2%) being the least represented. Conclusions and Relevance: Existing evaluations of LLMs mostly focus on accuracy of question answering for medical examinations, without consideration of real patient care data. Dimensions such as fairness, bias, and toxicity and deployment considerations received limited attention. Future evaluations should adopt standardized applications and metrics, use clinical data, and broaden focus to include a wider range of tasks and specialties.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.