ArticleFrontiers in digital health2025
Accuracy of ChatGPT-3.5, ChatGPT-4o, Copilot, Gemini, Claude, and Perplexity in advising on lumbosacral radicular pain against clinical practice guidelines: cross-sectional study.
Article in Frontiers in digital health, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to trial NCT07733752 (When AI Is the First Clinician), which is not on this map. Cited by 12 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
When AI Is the First Clinician: Impact of Pre-Visit AI Use on Presentation, Diagnostic Expectations, and Shared Decision-Making in Spine Physical Therapy
Who cites it
12 citing papers in PubMed.
- Evaluating the quality of artificial intelligence responses to psoriasis-related clinical and patient questions: a comparative study of ChatGPT, Gemini, and Microsoft Copilot.Proceedings (Baylor University. Medical Center) · 2026Article
- Large Language Models in Spine Surgery : A Narrative Review of Performance Paradox and Clinical Integration Challenges.Journal of Korean Neurosurgical Society · 2026Article
- Generative AI and Large Language Models in Rehabilitation: A Scoping Review.Life (Basel, Switzerland) · 2026Review
- Evaluation of GPT-5, a Large Language Model, in Replicating German Clinical Practice Guideline Recommendations in Oral Oncology: A Cross-Sectional Concordance Study.Journal of oral pathology & medicine : official publication of the International Association of Oral Pathologists and the American Academy of Oral Pathology · 2026Article
- Article
- Promise and pitfalls of AI chatbots in complex decision-making for thyroid nodules and papillary thyroid cancer.European thyroid journal · 2026Article
- Evaluating the efficacy and readability of advanced large language models in responding to patients' frequently asked questions about chronic rhinosinusitis: a comparative analysis.European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2026Article
- Article
- Article
- Performance of five free large language models in dental trauma: a 30-day longitudinal benchmark study.Frontiers in oral health · 2025Article
- Knowledge, use and perceptions of artificial intelligence Chatbots among Italian physiotherapists: an online cross-sectional survey.Frontiers in digital health · 2025Article
- Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
10 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Introduction: Artificial Intelligence (AI) chatbots, which generate human-like responses based on extensive data, are becoming important tools in healthcare by providing information on health conditions, treatments, and preventive measures, acting as virtual assistants. However, their performance in aligning with clinical practice guidelines (CPGs) for providing answers to complex clinical questions on lumbosacral radicular pain is still unclear. We aim to evaluate AI chatbots' performance against CPG recommendations for diagnosing and treating lumbosacral radicular pain. Methods: We performed a cross-sectional study to assess AI chatbots' responses against CPGs recommendations for diagnosing and treating lumbosacral radicular pain. Clinical questions based on these CPGs were posed to the latest versions (updated in 2024) of six AI chatbots: ChatGPT-3.5, ChatGPT-4o, Microsoft Copilot, Google Gemini, Claude, and Perplexity. The chatbots' responses were evaluated for (a) consistency of text responses using Plagiarism Checker X, (b) intra- and inter-rater reliability using Fleiss' Kappa, and (c) match rate with CPGs. Statistical analyses were performed with STATA/MP 16.1. Results: We found high variability in the text consistency of AI chatbot responses (median range 26%-68%). Intra-rater reliability ranged from "almost perfect" to "substantial," while inter-rater reliability varied from "almost perfect" to "moderate." Perplexity had the highest match rate at 67%, followed by Google Gemini at 63%, and Microsoft Copilot at 44%. ChatGPT-3.5, ChatGPT-4o, and Claude showed the lowest performance, each with a 33% match rate. Conclusions: Despite the variability in internal consistency and good intra- and inter-rater reliability, the AI Chatbots' recommendations often did not align with CPGs recommendations for diagnosing and treating lumbosacral radicular pain. Clinicians and patients should exercise caution when relying on these AI models, since one to two-thirds of the recommendations provided may be inappropriate or misleading according to specific chatbots.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.