Evidence map›Paper›PMID 40489764›Full record

ArticleJournal of medical Internet research2025

Large Language Models in Medical Diagnostics: Scoping Review With Bibliometric Analysis.

Hankun Su, Yuanyuan Sun, Ruiting Li, Aozhe Zhang, Yuemeng Yang, Fen Xiao, Zhiying Duan, Jingjing Chen, Qin Hu, Tianli Yang and 5 more

Registry-linked trialAbstract readScoping Review
In one paragraph

Article in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to trial NCT07179861 (A Prospective, Cross-Sectional, Vignette-Based Observational Study Comparing Clinical Decision-Making Performance of Pediatriciansand AI Models), which is not on this map. Cited by 24 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
24citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

NCT07179861 completednot on this map

A Prospective, Cross-Sectional, Vignette-Based Observational Study Comparing Clinical Decision-Making Performance of Pediatriciansand AI Models

TypeobservationalSponsorHaseki Training and Research HospitalRan2025 to 2025Enrolled30ConditionsArtificial Intelligence (AI) in Diagnosis, Decision Support Systems, Clinical, Clinical Decision-making, PediatricsArmsAI Suggestions (Anonymized 5-tool panel), Confidence Rating Task (1-10 Likert)
3 · Its place in the literature

Who cites it

24 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Review
  3. Article
  4. Review
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Review
  16. Article
  17. AI-Powered Clinical Decision Support in Dentistry: Comparative Evaluation of Large Language Models for Oral Medicine and Periodontal Diagnosis.Medical science monitor : international medical journal of experimental and clinical research · 2026
    Article
  18. Review
  19. Article
  20. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

15 authors.

Hankun Su *Department of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0009-0008-1748-1978
Yuanyuan Sun *Department of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0009-0005-0612-8800
Ruiting LiSchool of Biomedical Sciences and Engineering, South China University of Technology, Guangzhou, China.ORCID https://orcid.org/0009-0002-1889-0227
Aozhe ZhangXiangya School of Medicine, Central South University, Changsha, China.ORCID https://orcid.org/0009-0005-8854-8306
Yuemeng YangDepartment of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0009-0008-9035-2689
Fen XiaoDepartment of Metabolism and Endocrinology, Second Xiangya Hospital of Central South University, Changsha, China.ORCID https://orcid.org/0000-0001-7375-3482
Zhiying DuanDepartment of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0009-0007-2337-9497
Jingjing ChenDepartment of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0000-0002-1857-8498
Qin HuDepartment of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0009-0007-3065-9155
Tianli YangDepartment of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0009-0006-5860-5258
Bin XuDepartment of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0000-0001-6338-4230
Qiong ZhangDepartment of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0009-0001-0023-5286
Jing ZhaoDepartment of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0000-0002-3658-3274
Yanping LiDepartment of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0000-0002-4193-7321
Hui LiDepartment of Reproductive Medicine, Xiangya Hospital Central South University, Changsha, China.ORCID https://orcid.org/0000-0002-7516-7126

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundThe integration of large language models (LLMs) into medical diagnostics has garnered substantial attention due to their potential to enhance diagnostic accuracy, streamline clinical workflows, and address health care disparities. However, the rapid evolution of LLM research necessitates a comprehensive synthesis of their applications, challenges, and future directions.

objectiveThis scoping review aimed to provide an overview of the current state of research regarding the use of LLMs in medical diagnostics. The study sought to answer four primary subquestions, as follows: (1) Which LLMs are commonly used? (2) How are LLMs assessed in diagnosis? (3) What is the current performance of LLMs in diagnosing diseases? (4) Which medical domains are investigating the application of LLMs?

methodsThis scoping review was conducted according to the Joanna Briggs Institute Manual for Evidence Synthesis and adheres to the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews). Relevant literature was searched from the Web of Science, PubMed, Embase, IEEE Xplore, and ACM Digital Library databases from 2022 to 2025. Articles were screened and selected based on predefined inclusion and exclusion criteria. Bibliometric analysis was performed using VOSviewer to identify major research clusters and trends. Data extraction included details on LLM types, application domains, and performance metrics.

resultsThe field is rapidly expanding, with a surge in publications after 2023. GPT-4 and its variants dominated research (70/95, 74% of studies), followed by GPT-3.5 (34/95, 36%). Key applications included disease classification (text or image-based), medical question answering, and diagnostic content generation. LLMs demonstrated high accuracy in specialties like radiology, psychiatry, and neurology but exhibited biases in race, gender, and cost predictions. Ethical concerns, including privacy risks and model hallucination, alongside regulatory fragmentation, were critical barriers to clinical adoption.

conclusionsLLMs hold transformative potential for medical diagnostics but require rigorous validation, bias mitigation, and multimodal integration to address real-world complexities. Future research should prioritize explainable artificial intelligence frameworks, specialty-specific optimization, and international regulatory harmonization to ensure equitable and safe clinical deployment.

Indexed as

BibliometricsLanguageHumansLarge Language Modelsartificial intelligencebibliometric analysislarge language modelmedical diagnosisscoping review

Identifiers

PMID40489764
PMCPMC12186007

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.