ArticleEuropean radiology2025
Reproducibility of methodological radiomics score (METRICS): an intra- and inter-rater reliability study endorsed by EuSoMII.
Article in European radiology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
6 citing papers in PubMed, 1 synthesis or guideline pooled it.
- CT and MRI radiomics in cardiovascular risk prediction: a systematic review and meta-analysis by the EuSoMII Radiomics Auditing Group.European radiology · 2026Pooled it
- Methodological quality of cardiac CT and MRI radiomics studies assessed using METRICS and RQS by human readers and ChatGPT 5.1 Thinking.European radiology experimental · 2026Article
- MRI-Based Radiomics to Predict Response to Neoadjuvant Therapy in Locally Advanced Rectal Cancer: A Retrospective Study.Journal of personalized medicine · 2026Article
- Explanation and Elaboration with Examples for METRICS (METRICS-E3): an initiative from the EuSoMII Radiomics Auditing Group.Insights into imaging · 2025Article
- Quality appraisal of radiomics-based studies on chondrosarcoma using METhodological RadiomICs Score (METRICS) and Radiomics Quality Score (RQS).Insights into imaging · 2025Article
- A multimodal machine learning model for predicting postoperative worsening of FOGQ in Parkinson's disease following STN-DBS.Frontiers in neurologyArticle
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
17 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
objectivesTo investigate the intra- and inter-rater reliability of the total methodological radiomics score (METRICS) and its items through a multi-reader analysis. MATERIALS AND
methodsA total of 12 raters with different backgrounds and experience levels were recruited for the study. Based on their level of expertise, raters were randomly assigned to the following groups: two inter-rater reliability groups, and two intra-rater reliability groups, where each group included one group with and one group without a preliminary training session on the use of METRICS. Inter-rater reliability groups assessed all 34 papers, while intra-rater reliability groups completed the assessment of 17 papers twice within 21 days each time, and a "wash out" period of 60 days in between.
resultsInter-rater reliability was poor to moderate between raters of group 1 (without training; ICC = 0.393; 95% CI = 0.115-0.630; p = 0.002), and between raters of group 2 (with training; ICC = 0.433; 95% CI = 0.127-0.671; p = 0.002). The intra-rater analysis was excellent for raters 9 and 12, good to excellent for raters 8 and 10, moderate to excellent for rater 7, and poor to good for rater 11.
conclusionThe intra-rater reliability of the METRICS score was relatively good, while the inter-rater reliability was relatively low. This highlights the need for further efforts to achieve a common understanding of METRICS items, as well as resources consisting of explanations, elaborations, and examples to improve reproducibility and enhance their usability and robustness. KEY POINTS: Questions Guidelines and scoring tools are necessary to improve the quality of radiomics research; however, the application of these tools is challenging for less experienced raters. Findings Intra-rater reliability was high across all raters regardless of experience level or previous training, and inter-rater reliability was generally poor to moderate across raters. Clinical relevance Guidelines and scoring tools are necessary for proper reporting in radiomics research and for closing the gap between research and clinical implementation. There is a need for further resources offering explanations, elaborations, and examples to enhance the usability and robustness of these guidelines.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.