ArticleJournal of medical Internet research2026
Publicly Accessible Large Language Model Responses to Frequently Asked Questions About Spondylodiscitis: Preliminary Expert Evaluation.
Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
13 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Patients increasingly use large language models (LLMs) to obtain medical information, but the quality of LLM-generated information on complex spinal infections such as spondylodiscitis remains uncertain. Existing evaluations in spine surgery have mainly addressed degenerative conditions or surgical procedures, and disease-specific data for spondylodiscitis are limited. Objective: This preliminary study evaluated spine surgeons' ratings of single-turn LLM responses to 10 author-curated frequently asked questions (FAQs) about spondylodiscitis and compared answer sets generated from GPT-4, GPT-4o, and Google Gemini web interfaces under the authors' implemented prompting conditions. Methods: A pool of patient-oriented questions was generated through a chronological workflow including publicly available FAQ sources, PubMed-informed terminology review, Google Trends topic checking, and LLM-generated candidate questions. Duplicate and semantically overlapping questions were removed, the remaining questions were grouped into thematic categories, and 10 final FAQ-style prompts were synthesized. Each prompt was submitted once to the publicly accessible web interfaces of GPT-4, GPT-4o, and Google Gemini. The study was interpreted as an expert evaluation of the resulting answer sets. Seven blinded board-certified spine surgeons rated the responses using a 4-level rating system ranging from excellent to unsatisfactory and additionally assessed comprehensiveness, clarity, empathy, and appropriateness of length. Descriptive statistics and nonparametric comparisons were performed. Interrater reliability was assessed using the intraclass correlation coefficient. Results: Across all responses, 38.6% (81/210) were rated as excellent, 39% (82/210) as satisfactory with minimal clarification needed, 16.7% (35/210) as satisfactory with moderate clarification needed, and 5.7% (12/210) as unsatisfactory. The most common reason for necessary clarification was insufficient information (58/141, 41.1%), followed by language-related issues (21/141, 14.9%) and overly detailed responses (18/141, 12.8%). The complication-related question received the highest mean rating (3.4/5), whereas treatment- and prognosis-related questions received lower ratings (2.7/5 and 2.9/5). Median overall ratings did not differ significantly among the 3 evaluated LLMs. Spine surgeons reported a generally positive attitude toward artificial intelligence-supported patient information but expressed remaining uncertainty regarding reliability and direct patient-physician communication. Conclusions: In this preliminary expert evaluation, selected publicly accessible LLM web interfaces generated mostly satisfactory responses to author-curated spondylodiscitis FAQs. The findings reflect outputs produced at the time of access under the authors' implemented prompting conditions. LLMs may support patient education only with clinician oversight. Future research should explore advanced, domain-specific models to further improve the quality of communication between clinicians and patients.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.