ArticleBMC medical informatics and decision making2021
Explaining multivariate molecular diagnostic tests via Shapley values.
Article in BMC medical informatics and decision making, 2021. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 17 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
17 citing papers in PubMed.
- Predicting delayed graft function after kidney transplant: Do complex models help compared to standard statistics?World journal of nephrology · 2026Article
- Development and external validation of an online interpretable machine-learning model for predicting delirium risk in acute heart failure.The Journal of international medical research · 2026Article
- Decoding causal m6A: a bioinformatics roadmap for psychiatric disorders.Briefings in bioinformatics · 2026Review
- Fast multimodal imaging combined with machine learning identifying taurine as a potential marker for breast cancer margin assessment.NPJ digital medicine · 2025Article
- Innovations in early detection of chronic non-communicable diseases among adolescents through an easy-to-Use AutoML paradigm.Health care management science · 2025Observational
- Rigorous validation of machine learning in laboratory medicine: guidance toward quality improvement.Critical reviews in clinical laboratory sciences · 2025Review
- Predicting Visual Acuity after Retinal Vein Occlusion Anti-VEGF Treatment: Development and Validation of an Interpretable Machine Learning Model.Journal of medical systems · 2025Article
- Prediction of Mycobacterium tuberculosis cell wall permeability using machine learning methods.Molecular diversity · 2024Article
- Automated Machine Learning and Explainable AI (AutoML-XAI) for Metabolomics: Improving Cancer Diagnostics.Journal of the American Society for Mass Spectrometry · 2024Article
- Impossibility theorems for feature attribution.Proceedings of the National Academy of Sciences of the United States of America · 2024Article
- A machine learning-derived risk score to predict left ventricular diastolic dysfunction from clinical cardiovascular magnetic resonance imaging.Frontiers in cardiovascular medicine · 2024Article
- Automated machine learning and explainable AI (AutoML-XAI) for metabolomics: improving cancer diagnostics.bioRxiv : the preprint server for biology · 2023Article
- WindowSHAP: An efficient framework for explaining time-series classifiers based on Shapley values.Journal of biomedical informatics · 2023Article
- Ageing and cancer: a research gap to fill.Molecular oncology · 2022Review
- Development and Validation of an Insulin Resistance Model for a Population with Chronic Kidney Disease Using a Machine Learning Approach.Nutrients · 2022Article
- Real-world performance of blood-based proteomic profiling in first-line immunotherapy treatment in advanced stage non-small cell lung cancer.Journal for immunotherapy of cancer · 2021Article
- Construction and validation of risk prediction models for different subtypes of retinal vein occlusion.Advances in ophthalmology practice and researchArticle
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundMachine learning (ML) can be an effective tool to extract information from attribute-rich molecular datasets for the generation of molecular diagnostic tests. However, the way in which the resulting scores or classifications are produced from the input data may not be transparent. Algorithmic explainability or interpretability has become a focus of ML research. Shapley values, first introduced in game theory, can provide explanations of the result generated from a specific set of input data by a complex ML algorithm.
methodsFor a multivariate molecular diagnostic test in clinical use (the VeriStrat® test), we calculate and discuss the interpretation of exact Shapley values. We also employ some standard approximation techniques for Shapley value computation (local interpretable model-agnostic explanation (LIME) and Shapley Additive Explanations (SHAP) based methods) and compare the results with exact Shapley values.
resultsExact Shapley values calculated for data collected from a cohort of 256 patients showed that the relative importance of attributes for test classification varied by sample. While all eight features used in the VeriStrat® test contributed equally to classification for some samples, other samples showed more complex patterns of attribute importance for classification generation. Exact Shapley values and Shapley-based interaction metrics were able to provide interpretable classification explanations at the sample or patient level, while patient subgroups could be defined by comparing Shapley value profiles between patients. LIME and SHAP approximation approaches, even those seeking to include correlations between attributes, produced results that were quantitatively and, in some cases qualitatively, different from the exact Shapley values.
conclusionsShapley values can be used to determine the relative importance of input attributes to the result generated by a multivariate molecular diagnostic test for an individual sample or patient. Patient subgroups defined by Shapley value profiles may motivate translational research. However, correlations inherent in molecular data and the typically small ML training sets available for molecular diagnostic test development may cause some approximation methods to produce approximate Shapley values that differ both qualitatively and quantitatively from exact Shapley values. Hence, caution is advised when using approximate methods to evaluate Shapley explanations of the results of molecular diagnostic tests.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.