Evidence map›Paper›PMID 40305085›Full record

SynthesisJournal of medical Internet research2025

Accuracy of Large Language Models When Answering Clinical Research Questions: Systematic Review and Network Meta-Analysis.

Ling Wang, Jinglin Li, Boyang Zhuang, Shasha Huang, Meilin Fang, Cunze Wang, Wen Li, Mohan Zhang, Shurong Gong

Abstract readNetwork Meta-AnalysisSystematic Review
In one paragraph

Synthesis in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 32 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
32citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

32 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Blinded by the Bot: Benchmarking GPT and Gemini Against Human Authors in Otolaryngology Reviews.World journal of otorhinolaryngology - head and neck surgery · 2026
    Article
  9. AI in respiratory care: findings from the GOLD report.Journal of translational medicine · 2026
    Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Article
  16. Article
  17. Article
  18. Article
  19. Article
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Ling Wang *Fuzhou University Affiliated Provincial Hospital, Shengli Clinical Medical College, Fujian Medical University, Fuzhou, China.ORCID https://orcid.org/0000-0002-1970-0862
Jinglin Li *School of Pharmacy, Fujian Medical University, Fuzhou, China.ORCID https://orcid.org/0009-0004-4680-1548
Boyang Zhuang *Fujian Center For Drug Evaluation and Monitoring, Fuzhou, China.ORCID https://orcid.org/0009-0009-2162-7076
Shasha Huang *School of Pharmacy, Fujian University of Traditional Chinese Medicine, Fuzhou, China.ORCID https://orcid.org/0009-0003-3554-5828
Meilin Fang *School of Pharmacy, Fujian Medical University, Fuzhou, China.ORCID https://orcid.org/0009-0006-3499-9168
Cunze WangSchool of Pharmacy, Fujian Medical University, Fuzhou, China.ORCID https://orcid.org/0000-0002-7751-1242
Wen LiFuzhou University Affiliated Provincial Hospital, Shengli Clinical Medical College, Fujian Medical University, Fuzhou, China.ORCID https://orcid.org/0009-0005-0542-7489
Mohan ZhangSchool of Pharmacy, Fujian Medical University, Fuzhou, China.ORCID https://orcid.org/0009-0001-1786-0176
Shurong GongThe Third Department of Critical Care Medicine, Fuzhou University Affiliated Provincial Hospital, Shengli Clinical Medical College, Fujian Medical University, Fuzhou, Fujian, China.ORCID https://orcid.org/0000-0003-1746-8198

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundLarge language models (LLMs) have flourished and gradually become an important research and application direction in the medical field. However, due to the high degree of specialization, complexity, and specificity of medicine, which results in extremely high accuracy requirements, controversy remains about whether LLMs can be used in the medical field. More studies have evaluated the performance of various types of LLMs in medicine, but the conclusions are inconsistent.

objectiveThis study uses a network meta-analysis (NMA) to assess the accuracy of LLMs when answering clinical research questions to provide high-level evidence-based evidence for its future development and application in the medical field.

methodsIn this systematic review and NMA, we searched PubMed, Embase, Web of Science, and Scopus from inception until October 14, 2024. Studies on the accuracy of LLMs when answering clinical research questions were included and screened by reading published reports. The systematic review and NMA were conducted to compare the accuracy of different LLMs when answering clinical research questions, including objective questions, open-ended questions, top 1 diagnosis, top 3 diagnosis, top 5 diagnosis, and triage and classification. The NMA was performed using Bayesian frequency theory methods. Indirect intercomparisons between programs were performed using a grading scale. A larger surface under the cumulative ranking curve (SUCRA) value indicates a higher ranking of the corresponding LLM accuracy.

resultsThe systematic review and NMA examined 168 articles encompassing 35,896 questions and 3063 clinical cases. Of the 168 studies, 40 (23.8%) were considered to have a low risk of bias, 128 (76.2%) had a moderate risk, and none were rated as having a high risk. ChatGPT-4o (SUCRA=0.9207) demonstrated strong performance in terms of accuracy for objective questions, followed by Aeyeconsult (SUCRA=0.9187) and ChatGPT-4 (SUCRA=0.8087). ChatGPT-4 (SUCRA=0.8708) excelled at answering open-ended questions. In terms of accuracy for top 1 diagnosis and top 3 diagnosis of clinical cases, human experts (SUCRA=0.9001 and SUCRA=0.7126, respectively) ranked the highest, while Claude 3 Opus (SUCRA=0.9672) performed well at the top 5 diagnosis. Gemini (SUCRA=0.9649) had the highest rated SUCRA value for accuracy in the area of triage and classification.

conclusionsOur study indicates that ChatGPT-4o has an advantage when answering objective questions. For open-ended questions, ChatGPT-4 may be more credible. Humans are more accurate at the top 1 diagnosis and top 3 diagnosis. Claude 3 Opus performs better at the top 5 diagnosis, while for triage and classification, Gemini is more advantageous. This analysis offers valuable insights for clinicians and medical practitioners, empowering them to effectively leverage LLMs for improved decision-making in learning, diagnosis, and management of various clinical scenarios.

trial registrationPROSPERO CRD42024558245; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024558245.

Indexed as

Biomedical ResearchLarge Language ModelsBayes TheoremHumansaccuracyclinical research questionslarge language modelsLLMnetwork meta-analysisPRISMA

Identifiers

PMID40305085
PMCPMC12079073

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.