Evidence map›Paper›PMID 40954311›Full record

ReviewNature medicine2025

The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence.

Viknesh Sounderajah, Ahmad Guni, Xiaoxuan Liu, Gary S Collins, Alan Karthikesalingam, Sheraz R Markar, Robert M Golub, Alastair K Denniston, Shravya Shetty, David Moher and 4 more

Erratum issued 2 registry-linked trialsAbstract readReview
PubMed Publisher
In one paragraph

Review in Nature medicine, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. An erratum has been issued. It is linked to 2 registered trials, which are not on this map. Cited by 160 papers, 4 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
160citing papers in PubMed, 4 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

NCT07632859 completednot on this mapstarted 2026, after this paper: background citation

Diagnostic Accuracy of Two Large Language Models Against a Blinded Specialist Consensus Standard in Turkish Emergency Department Notes: A Retrospective Study of 600 Cases

TypeobservationalSponsorMarmara University Pendik Training and Research HospitalRan2026 to 2026Enrolled600ConditionsEmergency Medicine, Diagnostic Errors, Artificial Intelligence (AI) in Diagnosis
NCT07740122 recruitingnot on this mapstarted 2026, after this paper: background citation

FECAL-AI: Prospective Observational Validation of AI-Based Stool Image Analysis Against Quantitative Fecal Immunochemical Testing for Colorectal Neoplasia Risk Assessment

Typeobservational_patient_registrySponsorServicio de Salud Metropolitano Sur OrienteRan2026 to 2026Enrolled250ConditionsColorectal NeoplasmsArmsAI-Based Stool Image Analysis
3 · Its place in the literature

Who cites it

160 citing papers in PubMed, 4 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Pooled it
  4. Pooled it
  5. Review
  6. Article
  7. Accountability for large language models in health care.Bulletin of the World Health Organization · 2026
    Article
  8. Article
  9. Article
  10. Review
  11. AI In Leukemia Diagnostics: Complementing the Pathologist's Role.International journal of laboratory hematology · 2026
    Review
  12. Article
  13. Article
  14. Article
  15. Review
  16. Article
  17. Article
  18. Review
  19. Review
  20. Article

100 more citing papers are in PubMed but not listed here.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

14 authors.

Viknesh Sounderajah *Institute of Global Health Innovation, Imperial College London, London, UK.
Ahmad Guni *Institute of Global Health Innovation, Imperial College London, London, UK.
Xiaoxuan LiuBirmingham Health Partners Centre for Regulatory Science and Innovation, Birmingham, UK.ORCID http://orcid.org/0000-0002-1286-0038
Gary S CollinsCentre for Statistics in Medicine, Nuffield Department of Orthopaedics, Rheumatology and Musculoskeletal Sciences, University of Oxford, Oxford, UK.ORCID http://orcid.org/0000-0002-2772-2316
Alan KarthikesalingamGoogle Research, London, UK.ORCID http://orcid.org/0009-0000-4958-5976
Sheraz R MarkarDepartment of Surgery and Cancer, Imperial College London, London, UK.
Robert M GolubNorthwestern University Feinberg School of Medicine, Chicago, IL, USA.ORCID http://orcid.org/0009-0000-3270-0632
Alastair K DennistonBirmingham Health Partners Centre for Regulatory Science and Innovation, Birmingham, UK.ORCID http://orcid.org/0000-0001-7849-0087
Shravya ShettyGoogle Health, Palo Alto, CA, USA.ORCID http://orcid.org/0000-0003-3783-3172
David MoherOttawa Hospital Research Institute, Ottawa, Ontario, Canada.ORCID http://orcid.org/0000-0003-2434-4206
Patrick M BossuytDepartment of Epidemiology and Data Science, Amsterdam University Medical Centres, Duivendrecht, the Netherlands.
Ara DarziInstitute of Global Health Innovation, Imperial College London, London, UK.ORCID http://orcid.org/0000-0001-7815-7989
Hutan AshrafianInstitute of Global Health Innovation, Imperial College London, London, UK. h.ashrafian@imperial.ac.uk.ORCID http://orcid.org/0000-0003-1668-0672
STARD-AI Steering Committee

Funding

Cancer Research UK (CRUK) C49297/A27294
6 · The paper itself

Abstract

The Standards for Reporting Diagnostic Accuracy (STARD) 2015 statement facilitates transparent and complete reporting of diagnostic test accuracy studies. However, there are unique considerations associated with artificial intelligence (AI)-centered diagnostic test studies. The STARD-AI statement, which was developed through a multistage, multistakeholder process, provides a minimum set of criteria that allows for comprehensive reporting of AI-centered diagnostic test accuracy studies. The process involved a literature review, a scoping survey of international experts, and a patient and public involvement and engagement initiative, culminating in a modified Delphi consensus process involving over 240 international stakeholders and a consensus meeting. The checklist was subsequently finalized by the Steering Committee and includes 18 new or modified items in addition to the STARD 2015 checklist items. Authors are encouraged to provide descriptions of dataset practices, the AI index test and how it was evaluated, as well as considerations of algorithmic bias and fairness. The STARD-AI statement supports comprehensive and transparent reporting in all AI-centered diagnostic accuracy studies, and it can help key stakeholders to evaluate the biases, applicability and generalizability of study findings.

Indexed as

Artificial IntelligenceDiagnostic Tests, RoutineChecklistConsensusDelphi TechniqueGuidelines as TopicHumansResearch Design

Identifiers

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.