Evidence map›Paper›PMID 39958450›Full record

ArticleWorld journal of gastroenterology2025

Evaluating large language models as patient education tools for inflammatory bowel disease: A comparative study.

Yan Zhang, Xiao-Han Wan, Qing-Zhou Kong, Han Liu, Jun Liu, Jing Guo, Xiao-Yun Yang, Xiu-Li Zuo, Yan-Qing Li

Abstract readComparative StudyEvaluation Study
In one paragraph

Article in World journal of gastroenterology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 11 papers.

0numbers the graph read from it
0cells of the map it votes in
11citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

11 citing papers in PubMed.

  1. Article
  2. Review
  3. Review
  4. Article
  5. Article
  6. Observational
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Yan ZhangDepartment of Gastroenterology, Qilu Hospital of Shandong University, Jinan 250012, Shandong Province, China.
Xiao-Han WanDepartment of Gastroenterology, Qilu Hospital of Shandong University, Jinan 250012, Shandong Province, China.
Qing-Zhou KongDepartment of Gastroenterology, Qilu Hospital of Shandong University, Jinan 250012, Shandong Province, China.
Han LiuDepartment of Gastroenterology, Qilu Hospital of Shandong University, Jinan 250012, Shandong Province, China.
Jun LiuDepartment of Gastroenterology, Qilu Hospital of Shandong University, Jinan 250012, Shandong Province, China.
Jing GuoDepartment of Gastroenterology, Qilu Hospital of Shandong University, Jinan 250012, Shandong Province, China.
Xiao-Yun YangDepartment of Gastroenterology, Qilu Hospital of Shandong University, Jinan 250012, Shandong Province, China.
Xiu-Li ZuoDepartment of Gastroenterology, Qilu Hospital of Shandong University, Jinan 250012, Shandong Province, China.
Yan-Qing LiDepartment of Gastroenterology, Qilu Hospital of Shandong University, Jinan 250012, Shandong Province, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundInflammatory bowel disease (IBD) is a global health burden that affects millions of individuals worldwide, necessitating extensive patient education. Large language models (LLMs) hold promise for addressing patient information needs. However, LLM use to deliver accurate and comprehensible IBD-related medical information has yet to be thoroughly investigated.

aimTo assess the utility of three LLMs (ChatGPT-4.0, Claude-3-Opus, and Gemini-1.5-Pro) as a reference point for patients with IBD.

methodsIn this comparative study, two gastroenterology experts generated 15 IBD-related questions that reflected common patient concerns. These questions were used to evaluate the performance of the three LLMs. The answers provided by each model were independently assessed by three IBD-related medical experts using a Likert scale focusing on accuracy, comprehensibility, and correlation. Simultaneously, three patients were invited to evaluate the comprehensibility of their answers. Finally, a readability assessment was performed.

resultsOverall, each of the LLMs achieved satisfactory levels of accuracy, comprehensibility, and completeness when answering IBD-related questions, although their performance varies. All of the investigated models demonstrated strengths in providing basic disease information such as IBD definition as well as its common symptoms and diagnostic methods. Nevertheless, when dealing with more complex medical advice, such as medication side effects, dietary adjustments, and complication risks, the quality of answers was inconsistent between the LLMs. Notably, Claude-3-Opus generated answers with better readability than the other two models.

conclusionLLMs have the potential as educational tools for patients with IBD; however, there are discrepancies between the models. Further optimization and the development of specialized models are necessary to ensure the accuracy and safety of the information provided.

Indexed as

Inflammatory Bowel DiseasesPatient Education as TopicAdultComprehensionFemaleGastroenterologyHumansLarge Language ModelsMaleSurveys and QuestionnairesInflammatory bowel diseaseLarge language modelsMedical information accuracyPatient educationReadability assessment

Identifiers

PMID39958450
PMCPMC11752706

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.