Evidence map›Paper›PMID 41176565›Full record

ArticleEuropean archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery2026

Can LLMs simplify operative notes? A comparative analysis in otorhinolaryngology.

Ahmet Ufuk Kılıçtaş, Oğuz Gül, Bilgeşah Kılıçtaş, Esat Kaba, Basar Erdivanli

Abstract readComparative StudyEvaluation StudyComparative Study
PubMed Publisher
In one paragraph

Article in European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Ahmet Ufuk KılıçtaşDepartment of Otorhinolaryngology, Konya City Hospital, Konya, Türkiye, Turkey. aukilictas@icloud.com.ORCID http://orcid.org/0000-0002-2537-5277
Oğuz GülDepartment of Otorhinolaryngology, Akçaabat Haçkalı Baba State Hospital, Trabzon, Türkiye.ORCID http://orcid.org/0000-0003-4626-0820
Bilgeşah KılıçtaşDepartment of Medical Oncology, Faculty of Medicine, Necmettin Erbakan University, Konya, Türkiye.ORCID http://orcid.org/0000-0002-2485-7901
Esat KabaDepartment of Radiology, Recep Tayyip Erdogan University Training and Research Hospital, Rize, Türkiye.ORCID http://orcid.org/0000-0001-7464-988X
Basar ErdivanliDepartment of Anesthesiology, Recep Tayyip Erdogan University, Rize, Türkiye.ORCID http://orcid.org/0000-0002-3955-8242

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

introductionOperative notes play a critical role in documenting surgical procedures and supporting medical communication. However, due to their technical language, these documents are often complex and difficult to understand for patients, non-medical individuals, and even some healthcare professionals. Large Language Models (LLMs) offer a novel opportunity to simplify such documents and make them more accessible. This study aims to quantify how six LLMs simplify otolaryngology operative notes and to compare readability, clinical accuracy and clarity. MATERIALS AND

methodsIn this study, 39 fictional operative notes specific to otolaryngologic surgery were simplified using six LLMs (GPT-4, GPT-4o, Claude 3.7, Gemini 2.0, DeepSeek, and Microsoft Copilot). The outputs were analyzed using eight different readability metrics and evaluated by two expert physicians in terms of medical accuracy and comprehensibility. Correlation analyses were also conducted across clinical subgroups (rhinology, otology, head and neck surgery).

resultsClaude 3.7 produced the most complex outputs, whereas GPT-4o, Gemini, and DeepSeek generated the most readable texts. According to expert evaluations, GPT-4 achieved the highest scores for medical accuracy, while GPT-4o received the highest ratings for clarity. Model performance varied across clinical subgroups.

conclusionLLMs are effective tools for simplifying medical texts; however, model selection should consider the target audience and clinical context, and all outputs must be verified by medical experts. When used in a controlled and validated manner, LLMs may contribute significantly to a new era of health communication. LEVEL OF EVIDENCE: N/A.

Indexed as

Large Language ModelsMedical Records Systems, ComputerizedOtolaryngologyOtorhinolaryngologic Surgical ProceduresComprehensionData AccuracyDocumentationHumansLarge language modelsMedical communicationOperative notesReadabilitySimplification

Identifiers

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.