Evidence map›Paper›PMID 34238309›Full record

ArticleBMC medical informatics and decision making2021

Explaining multivariate molecular diagnostic tests via Shapley values.

Joanna Roder, Laura Maguire, Robert Georgantas, Heinrich Roder

Abstract read
In one paragraph

Article in BMC medical informatics and decision making, 2021. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 17 papers.

0numbers the graph read from it
0cells of the map it votes in
17citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

17 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
  4. Article
  5. Observational
  6. Review
  7. Article
  8. Article
  9. Article
  10. Impossibility theorems for feature attribution.Proceedings of the National Academy of Sciences of the United States of America · 2024
    Article
  11. Article
  12. Article
  13. Article
  14. Review
  15. Article
  16. Article
  17. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Joanna RoderBiodesix, Inc., 2970 Wilderness Place, Ste100, Boulder, CO, 80301, USA. joanna.roder@biodesix.com.
Laura MaguireBiodesix, Inc., 2970 Wilderness Place, Ste100, Boulder, CO, 80301, USA.
Robert GeorgantasBiodesix, Inc., 2970 Wilderness Place, Ste100, Boulder, CO, 80301, USA.
Heinrich RoderBiodesix, Inc., 2970 Wilderness Place, Ste100, Boulder, CO, 80301, USA.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundMachine learning (ML) can be an effective tool to extract information from attribute-rich molecular datasets for the generation of molecular diagnostic tests. However, the way in which the resulting scores or classifications are produced from the input data may not be transparent. Algorithmic explainability or interpretability has become a focus of ML research. Shapley values, first introduced in game theory, can provide explanations of the result generated from a specific set of input data by a complex ML algorithm.

methodsFor a multivariate molecular diagnostic test in clinical use (the VeriStrat® test), we calculate and discuss the interpretation of exact Shapley values. We also employ some standard approximation techniques for Shapley value computation (local interpretable model-agnostic explanation (LIME) and Shapley Additive Explanations (SHAP) based methods) and compare the results with exact Shapley values.

resultsExact Shapley values calculated for data collected from a cohort of 256 patients showed that the relative importance of attributes for test classification varied by sample. While all eight features used in the VeriStrat® test contributed equally to classification for some samples, other samples showed more complex patterns of attribute importance for classification generation. Exact Shapley values and Shapley-based interaction metrics were able to provide interpretable classification explanations at the sample or patient level, while patient subgroups could be defined by comparing Shapley value profiles between patients. LIME and SHAP approximation approaches, even those seeking to include correlations between attributes, produced results that were quantitatively and, in some cases qualitatively, different from the exact Shapley values.

conclusionsShapley values can be used to determine the relative importance of input attributes to the result generated by a multivariate molecular diagnostic test for an individual sample or patient. Patient subgroups defined by Shapley value profiles may motivate translational research. However, correlations inherent in molecular data and the typically small ML training sets available for molecular diagnostic test development may cause some approximation methods to produce approximate Shapley values that differ both qualitatively and quantitatively from exact Shapley values. Hence, caution is advised when using approximate methods to evaluate Shapley explanations of the results of molecular diagnostic tests.

Indexed as

Machine LearningPathology, MolecularAlgorithmsCohort StudiesHumansArtificial intelligenceExplainabilityInterpretabilityMachine learningMolecular diagnostic testShapley values

Identifiers

PMID34238309
PMCPMC8265031

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.