ArticleJournal of medical Internet research2025
Large Language Models in Medical Diagnostics: Scoping Review With Bibliometric Analysis.
Article in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to trial NCT07179861 (A Prospective, Cross-Sectional, Vignette-Based Observational Study Comparing Clinical Decision-Making Performance of Pediatriciansand AI Models), which is not on this map. Cited by 24 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
A Prospective, Cross-Sectional, Vignette-Based Observational Study Comparing Clinical Decision-Making Performance of Pediatriciansand AI Models
Who cites it
24 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Diagnostic accuracy of large language models for emergency department triage: a systematic review and meta-analysis.BMC emergency medicine · 2026Pooled it
- How to apply artificial intelligence (AI) to facilitate and enhance MASH trials.Hepatology international · 2026Review
- Attitudes Toward Large Language Models in Health Care and Preferences for Their Adoption and Oversight Among Health Care Professionals: Cross-Sectional Survey.Journal of medical Internet research · 2026Article
- Large Language Models in Spine Surgery: A Scoping Review of Clinical Efficacy, Technical Integration, and Ethical Paradigms.Global spine journal · 2026Review
- The Role of Large Language Models in Hormonal Contraception Consultation.Journal of clinical medicine · 2026Article
- Diagnostic Performance and Workup Efficiency of Large Language Models in Secondary Hypertension: A Blinded Comparative Study.Diagnostics (Basel, Switzerland) · 2026Article
- Symptom-Only Localization of Brainstem Ischemia Using Large Language Models Versus Neurologists in Diffusion-Weighted Imaging-Positive Cases: Retrospective Single-Center Study.JMIR formative research · 2026Article
- Informed Consent Disclosures and Minimum Requirements in AI Clinical Trials: Cross-Sectional Analysis.Journal of medical Internet research · 2026Article
- ChatGPT-5 vs oral medicine experts for rank-based differential diagnosis of oral lesions: a prospective, biopsy-validated comparison.Odontology · 2026Article
- Article
- Applications of DeepSeek in Medicine: Bibliometric Analysis and Scoping Review.Journal of medical Internet research · 2026Article
- Gender-Attributed Persona Prompts and the Diagnostic Accuracy of Proprietary and Open-Weight Large Language Models in Chagas Disease and Visceral Leishmaniasis: A Paired Experimental Study.Healthcare (Basel, Switzerland) · 2026Article
- Public Expectations for Food and Drug Administration Approval of AI-Based Clinical Decision Support Tools: Quantitative Study.JMIR AI · 2026Article
- Open-Source Large Language Models Distilled DeepSeek-R1 Pose Challenges for On-Premises Clinical Deployment in Medical Diagnosis: A Comparative Study of Performance.Journal of medical systems · 2026Article
- A Decade of Artificial Intelligence in Stroke Care (2015-2025): Trends, Clinical Translation, and the Precision Medicine Frontier-A Narrative Review.Journal of personalized medicine · 2026Review
- Exploring Nurses' Perspectives on the Use of Artificial Intelligence Chatbots for Mental Health Support: A Cross-Sectional Study in Greece.Nursing reports (Pavia, Italy) · 2026Article
- AI-Powered Clinical Decision Support in Dentistry: Comparative Evaluation of Large Language Models for Oral Medicine and Periodontal Diagnosis.Medical science monitor : international medical journal of experimental and clinical research · 2026Article
- Medical large language models and systems in the clinical application of spinal diseases: Current status, challenges, and future prospects.Journal of orthopaedic translation · 2026Review
- Human versus artificial intelligence in oral pathology diagnosis: a comparative study of ChatGPT, Grok, and MANUS.Scientific reports · 2026Article
- Artificial Intelligence in Haematologic Diagnostics: Current Applications and Future Perspectives.Acta haematologica · 2026Review
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
15 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundThe integration of large language models (LLMs) into medical diagnostics has garnered substantial attention due to their potential to enhance diagnostic accuracy, streamline clinical workflows, and address health care disparities. However, the rapid evolution of LLM research necessitates a comprehensive synthesis of their applications, challenges, and future directions.
objectiveThis scoping review aimed to provide an overview of the current state of research regarding the use of LLMs in medical diagnostics. The study sought to answer four primary subquestions, as follows: (1) Which LLMs are commonly used? (2) How are LLMs assessed in diagnosis? (3) What is the current performance of LLMs in diagnosing diseases? (4) Which medical domains are investigating the application of LLMs?
methodsThis scoping review was conducted according to the Joanna Briggs Institute Manual for Evidence Synthesis and adheres to the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews). Relevant literature was searched from the Web of Science, PubMed, Embase, IEEE Xplore, and ACM Digital Library databases from 2022 to 2025. Articles were screened and selected based on predefined inclusion and exclusion criteria. Bibliometric analysis was performed using VOSviewer to identify major research clusters and trends. Data extraction included details on LLM types, application domains, and performance metrics.
resultsThe field is rapidly expanding, with a surge in publications after 2023. GPT-4 and its variants dominated research (70/95, 74% of studies), followed by GPT-3.5 (34/95, 36%). Key applications included disease classification (text or image-based), medical question answering, and diagnostic content generation. LLMs demonstrated high accuracy in specialties like radiology, psychiatry, and neurology but exhibited biases in race, gender, and cost predictions. Ethical concerns, including privacy risks and model hallucination, alongside regulatory fragmentation, were critical barriers to clinical adoption.
conclusionsLLMs hold transformative potential for medical diagnostics but require rigorous validation, bias mitigation, and multimodal integration to address real-world complexities. Future research should prioritize explainable artificial intelligence frameworks, specialty-specific optimization, and international regulatory harmonization to ensure equitable and safe clinical deployment.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.