Evidence map›Paper›PMID 40657647›Full record

ArticleFrontiers in digital health2025

Accuracy of ChatGPT-3.5, ChatGPT-4o, Copilot, Gemini, Claude, and Perplexity in advising on lumbosacral radicular pain against clinical practice guidelines: cross-sectional study.

Giacomo Rossettini, Silvia Bargeri, Chad Cook, Stefania Guida, Alvisa Palese, Lia Rodeghiero, Paolo Pillastrini, Andrea Turolla, Greta Castellini, Silvia Gianola

Registry-linked trialAbstract read
In one paragraph

Article in Frontiers in digital health, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to trial NCT07733752 (When AI Is the First Clinician), which is not on this map. Cited by 12 papers.

0numbers the graph read from it
0cells of the map it votes in
12citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

NCT07733752 recruitingnot on this mapstarted 2026, after this paper: background citation

When AI Is the First Clinician: Impact of Pre-Visit AI Use on Presentation, Diagnostic Expectations, and Shared Decision-Making in Spine Physical Therapy

Typeobservational_patient_registrySponsorAssiut UniversityRan2026 to 2027Enrolled200ConditionsLow Back Pain, Neck Pain, RadiculopathyArmsre-visit AI symptom-checker use
3 · Its place in the literature

Who cites it

12 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
  4. Evaluation of GPT-5, a Large Language Model, in Replicating German Clinical Practice Guideline Recommendations in Oral Oncology: A Cross-Sectional Concordance Study.Journal of oral pathology & medicine : official publication of the International Association of Oral Pathologists and the American Academy of Oral Pathology · 2026
    Article
  5. Article
  6. Article
  7. Evaluating the efficacy and readability of advanced large language models in responding to patients' frequently asked questions about chronic rhinosinusitis: a comparative analysis.European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2026
    Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Giacomo Rossettini *School of Physiotherapy, University of Verona, Verona, Italy.
Silvia Bargeri *Unit of Clinical Epidemiology, IRCCS Istituto Ortopedico Galeazzi, Milan, Italy.
Chad CookDepartment of Orthopaedics, Duke University, Durham, NC, United States.
Stefania GuidaUnit of Clinical Epidemiology, IRCCS Istituto Ortopedico Galeazzi, Milan, Italy.
Alvisa PaleseDepartment of Medical Sciences, University of Udine, Udine, Italy.
Lia RodeghieroDepartment of Rehabilitation, Hospital of Merano (SABES-ASDAA), Teaching Hospital of Paracelsus Medical University (PMU), Merano-Meran, Italy.
Paolo PillastriniDepartment of Biomedical and Neuromotor Sciences (DIBINEM), Alma Mater University of Bologna, Bologna, Italy.
Andrea TurollaDepartment of Biomedical and Neuromotor Sciences (DIBINEM), Alma Mater University of Bologna, Bologna, Italy.
Greta Castellini *Unit of Clinical Epidemiology, IRCCS Istituto Ortopedico Galeazzi, Milan, Italy.
Silvia Gianola *Unit of Clinical Epidemiology, IRCCS Istituto Ortopedico Galeazzi, Milan, Italy.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Introduction: Artificial Intelligence (AI) chatbots, which generate human-like responses based on extensive data, are becoming important tools in healthcare by providing information on health conditions, treatments, and preventive measures, acting as virtual assistants. However, their performance in aligning with clinical practice guidelines (CPGs) for providing answers to complex clinical questions on lumbosacral radicular pain is still unclear. We aim to evaluate AI chatbots' performance against CPG recommendations for diagnosing and treating lumbosacral radicular pain. Methods: We performed a cross-sectional study to assess AI chatbots' responses against CPGs recommendations for diagnosing and treating lumbosacral radicular pain. Clinical questions based on these CPGs were posed to the latest versions (updated in 2024) of six AI chatbots: ChatGPT-3.5, ChatGPT-4o, Microsoft Copilot, Google Gemini, Claude, and Perplexity. The chatbots' responses were evaluated for (a) consistency of text responses using Plagiarism Checker X, (b) intra- and inter-rater reliability using Fleiss' Kappa, and (c) match rate with CPGs. Statistical analyses were performed with STATA/MP 16.1. Results: We found high variability in the text consistency of AI chatbot responses (median range 26%-68%). Intra-rater reliability ranged from "almost perfect" to "substantial," while inter-rater reliability varied from "almost perfect" to "moderate." Perplexity had the highest match rate at 67%, followed by Google Gemini at 63%, and Microsoft Copilot at 44%. ChatGPT-3.5, ChatGPT-4o, and Claude showed the lowest performance, each with a 33% match rate. Conclusions: Despite the variability in internal consistency and good intra- and inter-rater reliability, the AI Chatbots' recommendations often did not align with CPGs recommendations for diagnosing and treating lumbosacral radicular pain. Clinicians and patients should exercise caution when relying on these AI models, since one to two-thirds of the recommendations provided may be inappropriate or misleading according to specific chatbots.

Indexed as

artificial intelligencechatbotsChatGPTmachine learningmusculoskeletalnatural language processingorthopaedicsphysiotherapy

Identifiers

PMID40657647
PMCPMC12245906

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.