Evidence map›Paper›PMID 41795080›Full record

ArticleBMC oral health2026

Comparative assessment of quality, consistency, and reference accuracy of MIH-related clinical information generated by ChatGPT-4o and DeepSeek R1.

Zübeyde Uçar Gündoğar, Derya Sarıoğlu

Abstract readComparative Study
In one paragraph

Article in BMC oral health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Zübeyde Uçar GündoğarDepartment of Pediatric Dentistry, Faculty of Dentistry, Gaziantep University, Gaziantep, Türkiye. zubeydeucargundogar@hotmail.com.
Derya SarıoğluDepartment of Pediatric Dentistry, Faculty of Dentistry, Gaziantep University, Gaziantep, Türkiye.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundMolar incisor hypomineralization (MIH) is a clinically challenging developmental enamel defect that requires accurate diagnosis and nuanced management decisions. Although large language models (LLMs) are increasingly used as sources of clinical information in dentistry, the quality, consistency, and reference accuracy of their MIH-related explanations remain largely unexplored. To date, no study has systematically compared the clinical quality or citation accuracy of LLM-generated MIH information. This study comparatively evaluated how two widely used LLMs generate clinical information across repeated sessions.

methodsTwenty open-ended MIH questions were developed in accordance with current clinical guidelines and organized into four categories: diagnosis, etiology, treatment, and differential diagnosis. ChatGPT-4o and DeepSeek R1 were each prompted with all questions during three independent sessions (morning, afternoon, evening) on the same day, generating a total of 120 responses. All responses were anonymized and evaluated by 24 calibrated pediatric dentists using the five-point Global Quality Scale (GQS). References provided by the models were independently verified by two reviewers and categorized as real or fabricated. Statistical analyses included Shapiro–Wilk tests, paired t-tests, Wilcoxon signed-rank tests, repeated-measures ANOVA, Friedman tests, Holm-corrected post-hoc comparisons, and ICC(2,1), with significance set at p < 0.05.

resultsAcross all time points and all four MIH-related categories, DeepSeek R1 consistently achieved significantly higher GQS scores than ChatGPT-4o (all adjusted p < 0.005). Mean score differences ranged from + 0.36 to + 0.71, with the largest gap observed for etiology questions in the evening session. When overall scores were examined, DeepSeek R1 (4.44 ± 0.54) again outperformed ChatGPT-4o (3.99 ± 0.59) (t(23) = 11.83, p < 0.0001). Both models showed statistically significant but clinically small session-related variations, with acceptable reliability indicated by ICC values (0.72 for ChatGPT-4o; 0.77 for DeepSeek R1). Reference verification revealed notable fabrication rates in both models: ChatGPT-4o provided 46.5% real and 53.5% fake references, while DeepSeek R1 provided 34.2% real and 65.8% fake references.

conclusionsDeepSeek R1 and ChatGPT-4o each demonstrated distinct strengths in generating MIH-related clinical explanations, with DeepSeek providing more detailed and context-focused responses and ChatGPT-4o producing clearer, more structured overviews. Although the score differences were modest, they reflect meaningful variations when applied to a condition as diagnostically complex as MIH. Both models showed acceptable temporal stability; however, their substantial rates of fabricated references underscore the need for careful expert oversight. Overall, while LLMs may support early learning, patient communication, or preliminary clinical orientation, neither model currently meets the accuracy or citation standards required for autonomous clinical use in pediatric dentistry.

Indexed as

Large Language ModelsMolar HypomineralizationDiagnosis, DifferentialGenerative Artificial IntelligenceHumansReproducibility of ResultsArtificial intelligenceChatGPT-4oDeepSeek R1Large language modelsMolar incisor hypomineralizationPediatric dentistryQuality assessmentReference accuracy

Identifiers

PMID41795080
PMCPMC13081285

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.