Evidence map›Paper›PMID 41494139›Full record

ArticleJCO clinical cancer informatics2026

Artificial Intelligence-Assisted Error Detection in Complex Clinical Documentation: Leveraging Large Language Models to Enhance Patient Safety in Oncology.

Peter May, Sina Nokodian, Christoph Nuernbergk, Manuel Knauer, Maike Hefter, Aaron Becker von Rose, Florian Bassermann, Johannes Jung

Abstract read
In one paragraph

Article in JCO clinical cancer informatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Peter MayDepartment of Medicine III, School of Medicine and Health, Technical University of Munich, Munich, Germany.ORCID 0000-0002-9045-0595
Sina NokodianDepartment of Medicine III, School of Medicine and Health, Technical University of Munich, Munich, Germany.
Christoph NuernbergkDepartment of Medicine III, School of Medicine and Health, Technical University of Munich, Munich, Germany.ORCID 0009-0001-6001-0601
Manuel KnauerDepartment of Medicine III, School of Medicine and Health, Technical University of Munich, Munich, Germany.ORCID 0000-0002-3475-2396
Maike HefterDepartment of Medicine III, School of Medicine and Health, Technical University of Munich, Munich, Germany.
Aaron Becker von RoseDepartment of Emergency Medicine, School of Medicine and Health, Technical University of Munich, Munich, Germany.
Florian BassermannDepartment of Medicine III, School of Medicine and Health, Technical University of Munich, Munich, Germany.
Johannes JungDepartment of Medicine III, School of Medicine and Health, Technical University of Munich, Munich, Germany.ORCID 0000-0003-0137-7929

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

purposeIn high-risk specialties such as oncology, errors in clinical documentation can have severe consequences, highlighting a need for enhanced safety checks. We therefore aimed to evaluate the capability of frontier large language models (LLMs) to identify and correct errors in complex clinical documentation in oncology.

methodsWe conducted a two-phase evaluation. First, we assessed LLMs (GPT o4-mini and Gemini 2.5 Pro) on 1,000 synthetic clinical hematology/oncology vignettes with controlled errors, benchmarking against human expert data for error flag detection and sentence localization. Second, we evaluated advanced LLMs and a local LLM (Gemma 3 27B) against six clinicians in detecting single, predefined, and clinically relevant errors, such as wrong risk classifications or omission of critical medication within 90 synthetic discharge summaries from oncologic patients.

resultsLLMs outperformed human benchmark in error flag and sentence localization tasks, with Gemini 2.5 Pro achieving top accuracies of 0.928 and 0.915, respectively. Results were robust across subgroups and scalable, with simultaneous processing of up to 50 vignettes. Within complex discharge summaries, Gemini 2.5 Pro and GPT o4-mini-high identified 97.8% and 87.8% of injected errors, respectively, substantially exceeding the 47.8% average detection rate of human specialists. Gemma 3 27B detected 35.6% of errors. Analysis of error detection overlap revealed a synergistic potential for hybrid human-artificial intelligence (AI) systems.

conclusionFrontier LLMs exhibit superior error-detection capabilities and speed compared with both local models and human specialists, who are inherently time-constrained. Although synthetic data provide a controlled testbed, real-world evaluation across diverse errors and documentation styles remains critical. Advanced LLMs can serve as powerful assistants for clinical documentation reviews, substantially reducing the risk of oversight and clinician workload. Integrating LLM-driven error flagging into electronic health record workflows offers a promising strategy for enhancing documentation accuracy, treatment quality, and patient safety in oncology.

Indexed as

Artificial IntelligenceDocumentationMedical ErrorsMedical OncologyNeoplasmsPatient SafetyElectronic Health RecordsHumansLarge Language Models

Identifiers

PMID41494139
PMCPMC12794695

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.