Evidence map›Paper›PMID 42051679›Full record

ReviewJournal of clinical and experimental dentistry2026

Evaluation of ChatGPT's Accuracy, Repeatability, and Reasoning Ability in Prosthodontics Education: A Cross-Sectional Comparative Study with Prosthodontists.

Naila Perween, Punit Raj Singh Khurana, Anju Aggarwal, Aditya Chaudhary, Kartika Nitin Kumar, Sahba Hassan, Athulya R Sekhar

Abstract readReview
In one paragraph

Review in Journal of clinical and experimental dentistry, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Naila PerweenAssociate Professor, Department of Prosthodontics, ITS Dental College, Hospital and Research Centre, Greater Noida, Uttar Pradesh- 201310.
Punit Raj Singh KhuranaProfessor and Head, Department of Prosthodontics, ITS Dental College, Hospital and Research Centre, Greater Noida, Uttar Pradesh- 201310.
Anju AggarwalProfessor, Department of Prosthodontics, ITS Dental College, Hospital and Research Centre, Greater Noida, Uttar Pradesh- 201310.
Aditya ChaudharyProfessor, Department of Prosthodontics, ITS Dental College, Hospital and Research Centre, Greater Noida, Uttar Pradesh- 201310.
Kartika Nitin KumarProfessor, Department of Prosthodontics, ITS Dental College, Hospital and Research Centre, Greater Noida, Uttar Pradesh- 201310.
Sahba HassanAssistant Professor, Department of Prosthodontics, ITS Dental College, Hospital and Research Centre, Greater Noida, Uttar Pradesh- 201310.
Athulya R SekharTutor, Department of Prosthodontics, ESIC Dental College & Hospital, New Delhi- 110085.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: The integration of artificial intelligence (AI) tools like ChatGPT in dental education is increasing, yet their accuracy, reasoning quality, and reliability remain underexplored in specialized fields like prosthodontics. This study aimed to evaluate the performance of ChatGPT in answering prosthodontics-based questions by comparing its accuracy with that of experienced Prosthodontists, as well as assessing its repeatability and reasoning ability. Material and Methods: A cross-sectional observational study was conducted using 36 validated prosthodontics-based questions, categorized by difficulty (easy, medium, hard) and type (theoretical, clinical). Responses were obtained from a panel of Prosthodontists via Google Form and from ChatGPT 4-o mini version, twice daily for 15 days. Each group generated 1080 responses. Accuracy of ChatGPT's responses was compared with Prosthodontists' responses. ChatGPT's reliability was assessed using Intraclass Correlation Coefficient (ICC), Standard Error of Measurement (SEM), and Coefficient of Variation (CV). Five subject matter experts rated ChatGPT's reasoning quality on a 3-point Likert scale, and Pearson correlation was used to analyze the relationship between reasoning and accuracy. Results: Prosthodontists outperformed ChatGPT in overall accuracy (p < 0.05), with significant differences observed particularly for medium-difficulty and clinical questions. ChatGPT demonstrated fair reliability (ICC = 0.427), with SEM of 25.18 and CV of 61.7% indicating moderate variability. Reasoning analysis showed that 38.9% of ChatGPT's responses were rated strong, while 36.1% were rated poor. A significant positive correlation was found between reasoning quality and accuracy (r = 0.353, p = 0.035). Conclusions: ChatGPT demonstrates moderate ability in delivering accurate theoretical information but lacks consistency and clinical judgment. Its role should be limited to a supplementary aid in dental education, with expert oversight required to ensure accuracy and contextual relevance.

Identifiers

PMID42051679
PMCPMC13119599

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.