SynthesisJournal of medical Internet research2024
Examining the Role of Large Language Models in Orthopedics: Systematic Review.
Synthesis in Journal of medical Internet research, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 31 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
31 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Performance comparison and future perspectives of deep learning and classical machine learning in bone tumor applications: a systematic review (2019-2025).BMC medical informatics and decision making · 2026Pooled it
- Large language models show distinct multidimensional performance profiles and variable reproducibility in consensus-based bone stress injury.Journal of experimental orthopaedics · 2026Article
- Temporal Reproducibility of Fracture Interpretation in Forensic Radiography: A Multispecialty Comparison of Physicians and Vision Language Models Including Fracture Subtype Description.Tomography (Ann Arbor, Mich.) · 2026Article
- Can ChatGPT pass the polish national medical specialization examination in orthopedics and traumatology?Archives of orthopaedic and trauma surgery · 2026Article
- Psychological Risk Assessment in Plastic Surgery via a DeepSeek Large Language Model: A Retrospective Cohort Study.Aesthetic plastic surgery · 2026Article
- Diagnostic Performance of Large Language Models for Orthopedic-Related Rare Diseases and Their Impact on Physicians' Diagnostic Accuracy: 2-Stage Comparative Evaluation Study Based on the Chinese Rare Disease Catalog.Journal of medical Internet research · 2026Article
- Artificial intelligence advancements for orthopaedic clinical reasoning: longitudinal assessment of newer models (ChatGPT-5, Grok-3, Gemini 2.5 Flash) compared to clinicians.Archives of orthopaedic and trauma surgery · 2026Article
- Potential and Pitfalls of Multimodal Large Language Models in Cerebral Palsy Hip Surveillance: A Radiographic Interpretation Study Assessing Educational Utility.Journal of clinical medicine · 2026Article
- Clinical Evaluation of ChatGPT-5.3 Responses to Patient-Oriented Questions on Scoliosis: A Multidimensional Expert Analysis.Healthcare (Basel, Switzerland) · 2026Article
- ChatGPT in Orthopedic Trauma: Consistency, Accuracy, and Agreement With Textbook and Expert Opinion.Cureus · 2026Article
- Exploring the potential of artificial intelligence and machine learning in orthopaedic surgery.Bulletin of the Hospital for Joint Disease (2013) · 2026Review
- [Use of artificial intelligence for next-gen anamnesis and communication in orthopedics & trauma surgery : From chatbots to ambient intelligence].Orthopadie (Heidelberg, Germany) · 2026Review
- Automated Approaches of Text Simplification of Patient Education Materials: Scoping Review.Journal of medical Internet research · 2026Article
- Large Language Models for Clinical Narrative Processing: Methods, Applications, and Challenges.Methods and protocols · 2026Article
- The application of large language models in orthopedic postgraduate education: potentials, challenges, and future prospects.Journal of orthopaedic surgery and research · 2026Review
- Efficacy of Large Language Models for Screening of Systematic Reviews on Periprosthetic Joint Infection.Journal of clinical medicine · 2026Article
- Evaluation of ChatGPT-5 responses to patient-centered questions on stromal vascular fraction for knee osteoarthritis: fair to good quality and content.BMC musculoskeletal disorders · 2026Article
- Benchmarking large language models against human experts in rehabilitation medicine: a multidimensional evaluation.Journal of neuroengineering and rehabilitation · 2026Article
- Medical large language models and systems in the clinical application of spinal diseases: Current status, challenges, and future prospects.Journal of orthopaedic translation · 2026Review
- Gemini 1.5 Flash provides the most reliable content while ChatGPT-4o offers the highest readability for patient education on meniscal tears.Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
8 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundLarge language models (LLMs) can understand natural language and generate corresponding text, images, and even videos based on prompts, which holds great potential in medical scenarios. Orthopedics is a significant branch of medicine, and orthopedic diseases contribute to a significant socioeconomic burden, which could be alleviated by the application of LLMs. Several pioneers in orthopedics have conducted research on LLMs across various subspecialties to explore their performance in addressing different issues. However, there are currently few reviews and summaries of these studies, and a systematic summary of existing research is absent.
objectiveThe objective of this review was to comprehensively summarize research findings on the application of LLMs in the field of orthopedics and explore the potential opportunities and challenges.
methodsPubMed, Embase, and Cochrane Library databases were searched from January 1, 2014, to February 22, 2024, with the language limited to English. The terms, which included variants of "large language model," "generative artificial intelligence," "ChatGPT," and "orthopaedics," were divided into 2 categories: large language model and orthopedics. After completing the search, the study selection process was conducted according to the inclusion and exclusion criteria. The quality of the included studies was assessed using the revised Cochrane risk-of-bias tool for randomized trials and CONSORT-AI (Consolidated Standards of Reporting Trials-Artificial Intelligence) guidance. Data extraction and synthesis were conducted after the quality assessment.
resultsA total of 68 studies were selected. The application of LLMs in orthopedics involved the fields of clinical practice, education, research, and management. Of these 68 studies, 47 (69%) focused on clinical practice, 12 (18%) addressed orthopedic education, 8 (12%) were related to scientific research, and 1 (1%) pertained to the field of management. Of the 68 studies, only 8 (12%) recruited patients, and only 1 (1%) was a high-quality randomized controlled trial. ChatGPT was the most commonly mentioned LLM tool. There was considerable heterogeneity in the definition, measurement, and evaluation of the LLMs' performance across the different studies. For diagnostic tasks alone, the accuracy ranged from 55% to 93%. When performing disease classification tasks, ChatGPT with GPT-4's accuracy ranged from 2% to 100%. With regard to answering questions in orthopedic examinations, the scores ranged from 45% to 73.6% due to differences in models and test selections.
conclusionsLLMs cannot replace orthopedic professionals in the short term. However, using LLMs as copilots could be a potential approach to effectively enhance work efficiency at present. More high-quality clinical trials are needed in the future, aiming to identify optimal applications of LLMs and advance orthopedics toward higher efficiency and precision.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.