Evidence map›Paper›PMID 42066286›Full record

ArticleJournal of medical Internet research2026

Performance of AI Tools in Citing Retracted Literature : Content Analysis.

Sebastian Labenbacher, Maximilian Niederer, Sascha Hammer, Matthias Bader, Nikolaus Schreiber, Helmar Bornemann-Cimenti

Abstract readEvaluation Study
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Sebastian Labenbacher *Department of Anesthesiology and Intensive Care Medicine, Medical University of Graz, Auenbruggerplatz 5, Graz, 8036, Austria, 43 316-385-81843.ORCID http://orcid.org/0000-0002-6877-1528
Maximilian Niederer *Department of Anesthesiology and Intensive Care Medicine, Medical University of Graz, Auenbruggerplatz 5, Graz, 8036, Austria, 43 316-385-81843.ORCID http://orcid.org/0000-0001-5330-8138
Sascha HammerDepartment of Anesthesiology and Intensive Care Medicine, Medical University of Graz, Auenbruggerplatz 5, Graz, 8036, Austria, 43 316-385-81843.ORCID http://orcid.org/0009-0008-4062-2572
Matthias BaderDepartment of Anesthesiology and Intensive Care Medicine, Medical University of Graz, Auenbruggerplatz 5, Graz, 8036, Austria, 43 316-385-81843.ORCID http://orcid.org/0000-0002-4620-6042
Nikolaus SchreiberDepartment of Anesthesiology and Intensive Care Medicine, Medical University of Graz, Auenbruggerplatz 5, Graz, 8036, Austria, 43 316-385-81843.ORCID http://orcid.org/0000-0003-3691-8457
Helmar Bornemann-CimentiDepartment of Anesthesiology and Intensive Care Medicine, Medical University of Graz, Auenbruggerplatz 5, Graz, 8036, Austria, 43 316-385-81843.ORCID http://orcid.org/0000-0002-1201-3752

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Generative artificial intelligence (GenAI) tools are increasingly used in scientific research to support literature searches, evidence synthesis, and manuscript preparation. While these systems promise substantial efficiency gains, concerns have emerged regarding their reliability, particularly their tendency to cite inaccurate, fabricated, or retracted literature. The unrecognized inclusion of retracted studies poses a serious risk to research integrity and evidence-based decision-making. Whether commonly used GenAI tools can reliably detect, exclude, or transparently communicate the retraction status of scientific publications remains unclear. Objective: This study aimed to evaluate the ability of freely available GenAI tools to correctly handle retracted scientific articles during literature searches. Primary and secondary outcomes focused on accuracy, reliability, and consistency in recognizing retracted literature. Methods: In this pragmatic trial, nine widely used free-access GenAI tools (ChatGPT 4, ChatGPT 5, Claude, Gemini, Perplexity, Microsoft Copilot, SciSpace, ScienceOS, and Consensus) were evaluated. Each tool was asked five predefined, standardized questions addressing topic overview, article identification, article summarization, and explicit assessment of retraction status. Overall, 15 retracted articles (the 10 most cited and 5 most recently retracted as of May 23, 2025) were selected from the Retraction Watch database. All questions were repeated twice to assess intratool consistency. Responses were independently rated as correct or incorrect by 2 researchers. Descriptive statistics summarized performance, and comparisons between general-purpose and research-focused AI tools were conducted using descriptive statistics. Interreviewer agreement was assessed using Cohen kappa coefficient. Results: None of the evaluated AI tools consistently handled retracted articles correctly. No model achieved perfect accuracy across all question sets. ChatGPT 5 performed best, defined by the primary outcome of achieving fully correct responses to all five predefined tasks (5/5) for the highest number of retracted articles, correctly answering all five questions for 8 of 15 articles (53.3%). Research-focused tools (SciSpace, ScienceOS, and Consensus) failed to produce a single fully correct response set. Retracted articles were frequently included in topic overviews without warning, with error rates exceeding 40% in several tools. When specifically asked about retraction status, most systems failed to provide correct or complete information. OpenEvidence only reported data for a subset of our retracted articles as it is only used in health care literature. It demonstrated strong performance in topic overviews but low accuracy in identifying retracted articles. Conclusions: Freely available GenAI tools are currently not able to detect, exclude, or appropriately flag retracted scientific literature. The widespread and confident reproduction of retracted studies represents a substantial threat to research integrity, particularly in medical and evidence-based fields. Until retraction-aware verification mechanisms are systematically integrated, independent source checking remains essential when using AI-assisted literature tools.

Indexed as

Generative Artificial IntelligenceRetraction of Publication as TopicArtificial IntelligenceReproducibility of ResultsAIartificial intelligencedata accuracyethicsevidence-based Practiceretraction of publicationretractionsscientific misconduct

Identifiers

PMID42066286
PMCPMC13134821

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.