ReviewBMC medical informatics and decision making2024
Qualitative metrics from the biomedical literature for evaluating large language models in clinical decision-making: a narrative review.
Review in BMC medical informatics and decision making, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 19 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
19 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Large language models in real-world clinical workflows: a systematic review of applications and implementation.Frontiers in digital health · 2025Pooled it
- Verification of the impact of differences between objective and subjective evaluation methods on the interpretation of artificial intelligence systems generating mammogram reports using vision-language model.Japanese journal of radiology · 2026Article
- Evaluation of large language model responses to patient questions on oral anticoagulant therapy: a comparative expert assessment.Exploratory research in clinical and social pharmacy · 2026Article
- Comments on "A Comparative Study on the Use of DeepSeek-R1 and ChatGPT-4.5 in Different Aspects of Plastic Surgery".Aesthetic plastic surgery · 2026Article
- Statistical Analysis of the Performance of Open-Access Language Models: A Tool in the Clinical Laboratory.EJIFCC · 2026Article
- A S.C.O.R.E. framework for evaluating open-ended responses from large language models in healthcare.Cell reports. Medicine · 2026Article
- Developing a Quality Evaluation Index System for Health Conversational Artificial Intelligence: Mixed Methods Study.Journal of medical Internet research · 2026Article
- Quality assessment of large language model-generated prior authorization letters in nephrology.Frontiers in digital health · 2026Article
- Intelligence without intuition: a mixed-methods pilot study on reasoning models in musculoskeletal physiotherapy for low-back pain.Frontiers in digital health · 2026Article
- Continuous Glucose Monitoring Data Analysis 2.0: Functional Data Pattern Recognition and Artificial Intelligence Applications.Journal of diabetes science and technology · 2025Article
- Large Language Models in Medical Diagnostics: Scoping Review With Bibliometric Analysis.Journal of medical Internet research · 2025Article
- Advancing large language models as patient education tools for inflammatory bowel disease.World journal of gastroenterology · 2025Article
- Iterative refinement and goal articulation to optimize large language models for clinical information extraction.NPJ digital medicine · 2025Article
- Prompts, privacy, and personalized learning: integrating AI into nursing education-a qualitative study.BMC nursing · 2025Article
- Prompts to Table: Specification and Iterative Refinement for Clinical Information Extraction with Large Language Models.medRxiv : the preprint server for health sciences · 2025Article
- Demographic and Physical Determinants of Unhealthy Food Consumption in Polish Long-Term Care Facilities.Nutrients · 2025Article
- Population health management fit lifecycles in analytics.Frontiers in artificial intelligence · 2025Article
- Effectiveness of ChatGPT to provide esophageal cancer information: A SERVQUAL-based analysis.Digital healthArticle
- Review
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
9 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundThe large language models (LLMs), most notably ChatGPT, released since November 30, 2022, have prompted shifting attention to their use in medicine, particularly for supporting clinical decision-making. However, there is little consensus in the medical community on how LLM performance in clinical contexts should be evaluated.
methodsWe performed a literature review of PubMed to identify publications between December 1, 2022, and April 1, 2024, that discussed assessments of LLM-generated diagnoses or treatment plans.
resultsWe selected 108 relevant articles from PubMed for analysis. The most frequently used LLMs were GPT-3.5, GPT-4, Bard, LLaMa/Alpaca-based models, and Bing Chat. The five most frequently used criteria for scoring LLM outputs were "accuracy", "completeness", "appropriateness", "insight", and "consistency".
conclusionsThe most frequently used criteria for defining high-quality LLMs have been consistently selected by researchers over the past 1.5 years. We identified a high degree of variation in how studies reported their findings and assessed LLM performance. Standardized reporting of qualitative evaluation metrics that assess the quality of LLM outputs can be developed to facilitate research studies on LLMs in healthcare.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.