Evidence map›Paper›PMID 41909018›Full record

ArticleIndian journal of orthopaedics2026

Evaluating ChatGPT-5's Performance in Answering Common Patient Questions About Femoroacetabular Impingement and Hip Arthroscopy.

Maximilian Voss, Hannah Jaeger, Mikhail Salzmann, Robert Prill, Timoty Osterberger, Ingo J Banke, Nikolai Ramadanov

Abstract read
In one paragraph

Article in Indian journal of orthopaedics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. What the papers say.Journal of hip preservation surgery · 2026
    Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Maximilian VossCenter of Orthopaedics and Traumatology, Brandenburg Medical School, University Hospital Brandenburg/Havel, Brandenburg an der Havel, Germany.
Hannah JaegerCenter of Orthopaedics and Traumatology, Brandenburg Medical School, University Hospital Brandenburg/Havel, Brandenburg an der Havel, Germany.
Mikhail SalzmannCenter of Orthopaedics and Traumatology, Brandenburg Medical School, University Hospital Brandenburg/Havel, Brandenburg an der Havel, Germany.
Robert PrillCenter of Orthopaedics and Traumatology, Brandenburg Medical School, University Hospital Brandenburg/Havel, Brandenburg an der Havel, Germany.
Timoty OsterbergerClinic of Orthopaedics and Sports Orthopaedics, School of Medicine and Health, TUM University Hospital, Technical University of Munich, Munich, Germany.
Ingo J BankeClinic of Orthopaedics and Sports Orthopaedics, School of Medicine and Health, TUM University Hospital, Technical University of Munich, Munich, Germany.
Nikolai RamadanovCenter of Orthopaedics and Traumatology, Brandenburg Medical School, University Hospital Brandenburg/Havel, Brandenburg an der Havel, Germany.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Hip arthroscopy (HAS) is widely used to treat femoroacetabular impingement syndrome (FAIS), and many patients rely on online resources for medical information. Large language models (LLMs) such as ChatGPT have shown potential as supplementary educational tools in orthopedics; however, existing evaluations are limited to earlier model generations with variable accuracy and completeness. This study aimed to evaluate the accuracy, clarity, relevance, and completeness of ChatGPT-5 responses to common patient questions regarding FAIS and HAS. Methods: ChatGPT-5 was used to generate 25 frequently asked patient questions and corresponding answers related to hip preservation. Two fellowship-trained hip preservation surgeons independently evaluated each response using a five-point Likert scale across four predefined domains: relevance, accuracy, clarity, and completeness. Descriptive statistics were calculated as mean ± standard deviation for each domain. Inter-rater reliability was assessed using a two-way random-effects intraclass correlation coefficient with absolute agreement (ICC [2, 1]) and complemented by exact agreement percentages. Results: All responses received excellent scores, with mean values ranging from 4.84 ± 0.27 (completeness) to 5.00 ± 0.00 (relevance). Accuracy (4.97 ± 0.08) and clarity (4.91 ± 0.17) were near-perfect. ICC values demonstrated moderate to excellent agreement (0.70-0.81), complemented by high exact agreement rates (84-100%). No answer contained factually incorrect, misleading, or unsafe information. Minor reductions in completeness were attributable to occasional brevity rather than substantive omissions. Conclusion: ChatGPT-5 generated highly accurate, clear, and clinically appropriate patient-oriented explanations regarding FAIS and HAS, showing clear improvement compared with earlier ChatGPT versions. Although ChatGPT-5 represents a marked advancement in AI-based patient education, its use should be regarded as a complementary educational tool rather than a replacement for professional orthopedic counseling.

Indexed as

Artificial intelligenceChatGPT-5Femoroacetabular impingementHip arthroscopyLarge language models

Identifiers

PMID41909018
PMCPMC13031688

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.