Evidence map›Paper›PMID 42500068›Full record

ArticleWorld journal of otorhinolaryngology - head and neck surgery2026

Blinded by the Bot: Benchmarking GPT and Gemini Against Human Authors in Otolaryngology Reviews.

Sholem Hack, Rebecca Attal, Letizia Nitro, Cecilia Rosso, Anastasia Urbanelli, Antonio Mario Bulfamante, Luigi Angelo Vaira, Omar G Ahmed, Masayoshi Takashima

Abstract read
In one paragraph

Article in World journal of otorhinolaryngology - head and neck surgery, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Sholem HackCity St. George's University of London, School of Medicine Program Delivered by University of Nicosia at the Chaim Sheba Medical Center Ramat Gan Israel.ORCID https://orcid.org/0009-0001-0651-6994
Rebecca AttalCity St. George's University of London, School of Medicine Program Delivered by University of Nicosia at the Chaim Sheba Medical Center Ramat Gan Israel.
Letizia NitroOtolaryngology Unit, Santi Paolo E Carlo Hospital, Department of Health Sciences Università Degli Studi Di Milano Milan Italy.
Cecilia RossoOtolaryngology Unit, Santi Paolo E Carlo Hospital, Department of Health Sciences Università Degli Studi Di Milano Milan Italy.
Anastasia UrbanelliDepartment of Surgical Sciences, Otorhinolaryngology Unit University of Turin Turin Italy.
Antonio Mario BulfamanteOtolaryngology Unit, Department of Pediatric Surgery "Vittore Buzzi" Children Hospital Milan Italy.
Luigi Angelo VairaMaxillofacial Surgery Operative Unit, Department of Medicine Surgery and Pharmacy University of Sassari Sassari Italy.
Omar G AhmedDepartment of Otolaryngology-Head and Neck Surgery Houston Methodist Hospital Houston Texas USA.
Masayoshi TakashimaDepartment of Otolaryngology-Head and Neck Surgery Houston Methodist Hospital Houston Texas USA.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objective: To compare the quality of scientific review articles generated by two artificial intelligence systems, ChatGPT and Gemini, with those written by human authors in the field of otolaryngology. Methods: Two otolaryngology topics, chronic rhinosinusitis and infantile subglottic hemangioma, were selected. For each topic, four AI-generated reviews (GPT-4.0 and Gemini 2.0; narrative and PRISMA-style) and one human-authored peer-reviewed review were included, yielding a total of 10 manuscripts (8 AI-generated, 2 human-authored). A blinded panel of seven board-certified otolaryngologists evaluated all manuscripts using a 5-point Likert scale across seven domains: scientific accuracy, depth of content, citation quality, structure and organization, readability and tone, critical insight, and overall scientific quality. Group comparisons were performed using linear mixed-effects models with random intercepts for reviewer and manuscript. Interrater reliability was assessed using Shrout-Fleiss intraclass correlation coefficients (ICC). Manual verification of AI-generated references was conducted to assess citation accuracy and fabrication. Results: Human-authored manuscripts received the highest ratings across all domains (overall quality 4.50 ± 0.76). GPT-4.0 demonstrated moderate performance (2.71 ± 1.46 overall), while Gemini 2.0 scored lowest (2.14 ± 1.01). Mixed-effects modeling demonstrated significant group differences across all domains ( Conclusion: GPT generated fluent, stylistically strong reviews but remained significantly inferior to human-authored manuscripts in analytical depth and citation integrity. Gemini 2.0 underperformed across all domains and demonstrated a substantial rate of fabricated citations. As large language models become integrated into academic workflows, transparent disclosure, structured fact-checking, and human oversight remain essential to safeguard scientific reliability.

Indexed as

citation integritygenerative artificial intelligenceGPT‐4.0otolaryngologypeer reviewscientific writing

Identifiers

PMID42500068
PMCPMC13398715

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.