Evidence map›Paper›PMID 41827942›Full record

ReviewDiagnostics (Basel, Switzerland)2026

TRIAGE: Trustworthy Reporting and Assessment for Clinical Gain and Effectiveness of AI Models.

Farzaneh Fazilati, Mohammad Zakaria Rajabi, Nima Alihosseini, Mohaddeseh Esmaeili Farsani, Seyed Hasan Sandid, Shadi Zamani, Mehrshad Alirezaei Farahani, Fateme Biriaei, Fateme Sadeghipour, Mohammad Taha Mirshamsi and 2 more

Abstract readReview
In one paragraph

Review in Diagnostics (Basel, Switzerland), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Review
  2. Advances in Artificial Intelligence for Wrist Joint Injury Diagnosis.International journal of medical sciences · 2026
    Review
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

12 authors.

Farzaneh FazilatiDepartment of Biomedical Engineering, Faculty of Engineering, University of Isfahan, Isfahan 8174673441, Iran.
Mohammad Zakaria RajabiDepartment of Biomedical Engineering, School of Medicine, Tehran University of Medical Science, Tehran 88989487, Iran.
Nima AlihosseiniDepartment of Biomedical Engineering, Faculty of Engineering, University of Isfahan, Isfahan 8174673441, Iran.ORCID 0009-0008-5727-8907
Mohaddeseh Esmaeili FarsaniDepartment of Biomedical Engineering, Faculty of Engineering, University of Isfahan, Isfahan 8174673441, Iran.
Seyed Hasan SandidDepartment of Biomedical Engineering, Faculty of Engineering, University of Isfahan, Isfahan 8174673441, Iran.
Shadi ZamaniDepartment of Biomedical Engineering, Faculty of Engineering, University of Isfahan, Isfahan 8174673441, Iran.
Mehrshad Alirezaei FarahaniDepartment of Biomedical Engineering, Faculty of Engineering, University of Isfahan, Isfahan 8174673441, Iran.
Fateme BiriaeiDepartment of Biomedical Engineering, Faculty of Engineering, University of Isfahan, Isfahan 8174673441, Iran.
Fateme SadeghipourDepartment of Biomedical Engineering, Faculty of Engineering, University of Isfahan, Isfahan 8174673441, Iran.
Mohammad Taha MirshamsiDepartment of Biomedical Engineering, Faculty of Engineering, University of Isfahan, Isfahan 8174673441, Iran.
Mottahareh FahamiDepartment of Biomedical Engineering, Faculty of Engineering, University of Isfahan, Isfahan 8174673441, Iran.
Hamid Reza MaratebDepartment of Biomedical Engineering, Faculty of Engineering, University of Isfahan, Isfahan 8174673441, Iran.ORCID 0000-0003-4408-2397

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Machine learning (ML), including deep learning, kernel-based classifiers, and ensemble methods, is increasingly used to support clinical diagnosis in medical imaging, biosignal interpretation, and electronic health record (EHR)-based decision support. Despite rapid progress, many diagnostic AI studies still rely on limited retrospective evaluation and single summary measures (e.g., accuracy or AUC), creating a gap between reported model performance and evidence required for safe clinical adoption. This review proposes TRIAGE, a clinically grounded evaluation framework designed to organize diagnostic AI testing as an evidence pipeline aligned with real clinical use cases (screening, triage, second reading, and confirmatory testing). We summarize core discrimination metrics derived from the confusion matrix (sensitivity, specificity, predictive values, likelihood ratios, diagnostic odds ratio, and F-scores) and highlight the importance of prevalence and spectrum effects for interpreting predictive value and clinical workload. We further review evaluation strategies for multi-class and multi-label diagnostic tasks using appropriate aggregation methods (micro, macro, and weighted averaging) and set-based measures such as Hamming loss, exact match ratio, and Jaccard/IoU. Because diagnostic deployment is threshold-dependent, we integrate representation curves (ROC, precision-recall, lift, and cumulative gain) with calibration assessment and clinical utility analysis, including calibration slope, Brier score, and decision-curve analysis. We also address robustness and fairness evaluation, leakage-resistant validation designs (patient-grouped splits, stratified and temporal validation, and external validation), computational constraints relevant to deployment (latency, throughput, and energy use), and statistically sound model comparison with multiplicity control. A structured TRIAGE checklist table summarizing the evaluation parameters described in this review is provided in the main text to support reproducible and clinically interpretable reporting.

Indexed as

cross-validationmachine learning evaluationperformance metricsrepresentation curvesrobustnessstatistical significance testing

Identifiers

PMID41827942
PMCPMC12984829

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.