Evidence map›Paper›PMID 41387473›Full record

ArticleScientific reports2025

Multidimensional assessment of large language model responses to patient questions on gestational diabetes mellitus.

Betul Yigit Yalcın, Ummu Mutlu, Ayse Merve Ok, Ozge Telci Caklili, Hulya Hacisahinogullari, Gulsah Yenidunya Yalin, Ozlem Soyluk Selcukbiricik, Ayse Kubat Uzum, Kubilay Karsidag

Abstract read
In one paragraph

Article in Scientific reports, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Betul Yigit YalcınDepartment of Endocrinology, Faculty of Medicine, Istanbul University, Istanbul, Turkey. byigit667@gmail.com.ORCID 0000-0002-5533-330X
Ummu MutluDepartment of Endocrinology, Faculty of Medicine, Istanbul University, Istanbul, Turkey.ORCID 0000-0002-5259-7326
Ayse Merve OkDepartment of Endocrinology, Kocaeli City Hospital, Kocaeli, Turkey.ORCID 0000-0002-1074-1801
Ozge Telci CakliliDepartment of Endocrinology, Kocaeli City Hospital, Kocaeli, Turkey.ORCID 0000-0001-7566-5427
Hulya HacisahinogullariDepartment of Endocrinology, Faculty of Medicine, Istanbul University, Istanbul, Turkey.ORCID 0000-0001-9989-6473
Gulsah Yenidunya YalinDepartment of Endocrinology, Faculty of Medicine, Istanbul University, Istanbul, Turkey.ORCID 0000-0002-9013-5237
Ozlem Soyluk SelcukbiricikDepartment of Endocrinology, Faculty of Medicine, Istanbul University, Istanbul, Turkey.ORCID 0000-0003-0732-4764
Ayse Kubat UzumDepartment of Endocrinology, Faculty of Medicine, Istanbul University, Istanbul, Turkey.ORCID 0000-0003-0478-1193
Kubilay KarsidagDepartment of Endocrinology, Faculty of Medicine, Istanbul University, Istanbul, Turkey.ORCID 0000-0002-9332-1262

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Gestational diabetes mellitus (GDM) is a prevalent condition requiring accurate patient education, yet the reliability and readability of large language models (LLMs) in this context remain uncertain. This study evaluated the performance of four LLMs-ChatGPT-4o, Gemini 2.5 Pro, Grok 3.0, and DeepSeek R-1-using 25 patient-oriented questions derived from clinical scenarios. Seven endocrinologists independently rated the responses with the modified DISCERN (mDISCERN) instrument and the Global Quality Score (GQS). Readability was analyzed using the Flesch Reading Ease (FRES), Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), Coleman-Liau Index (CLI), and Simple Measure of Gobbledygook (SMOG), while lexical diversity was assessed through type-token ratio (TTR). Grok and Gemini obtained the highest mDISCERN and GQS scores, whereas ChatGPT performed significantly lower (p < 0.05). DeepSeek generated the most readable outputs, while Grok provided the longest and most complex responses. All models scored below the FRES threshold of 60 recommended for lay audiences. Response length showed strong positive correlations with mDISCERN and GQS, while TTR was inversely related to quality but positively associated with readability. These findings highlight variability among LLMs in GDM education and emphasize the need for model-specific improvements to ensure reliable patient-facing health information.

Indexed as

Diabetes, GestationalLanguagePatient Education as TopicAdultComprehensionFemaleHumansLarge Language ModelsPregnancySurveys and QuestionnairesArtificial intelligenceGestational diabetes mellitusLarge language modelsPatient educationReadability

Identifiers

PMID41387473
PMCPMC12705660

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.