Evidence map›Paper›PMID 36738712›Full record

ArticleComputers in biology and medicine2023

Explainable artificial intelligence model for identifying COVID-19 gene biomarkers.

Fatma Hilal Yagin, İpek Balikci Cicek, Abedalrhman Alkhateeb, Burak Yagin, Cemil Colak, Mohammad Azzeh, Sami Akbulut

Open access · greenAbstract read
In one paragraph

Article in Computers in biology and medicine, 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 38 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
38citing papers in PubMed, 1 pooled it
19.5field-weighted citation impact, top 1% of its field
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

38 citing papers in PubMed, 1 synthesis or guideline pooled it, 85 citations in OpenAlex.

  1. Pooled it
  2. Review
  3. Article
  4. Review
  5. Article
  6. Review
  7. Review
  8. Article
  9. Article
  10. Article
  11. Review
  12. Article
  13. Article
  14. Review
  15. Article
  16. Article
  17. Review
  18. Article
  19. Review
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors at 2 institutions in 2 countries.

Fatma Hilal YaginDepartment of Biostatistics and Medical Informatics, Faculty of Medicine, Inonu University, 44280, Malatya, Turkey. Electronic address: hilal.yagin@inonu.edu.tr.
İpek Balikci CicekDepartment of Biostatistics and Medical Informatics, Faculty of Medicine, Inonu University, 44280, Malatya, Turkey. Electronic address: ipek.balikci@inonu.edu.tr.
Abedalrhman AlkhateebSoftware Engineering Department, King Hussein School for Computing Sciences, Amman, Jordan. Electronic address: a.lkhateeb@psut.edu.jo.
Burak YaginDepartment of Biostatistics and Medical Informatics, Faculty of Medicine, Inonu University, 44280, Malatya, Turkey. Electronic address: burak.yagin@inonu.edu.tr.
Cemil ColakDepartment of Biostatistics and Medical Informatics, Faculty of Medicine, Inonu University, 44280, Malatya, Turkey. Electronic address: cemil.colak@inonu.edu.tr.
Mohammad AzzehData Science Department, King Hussein School for Computing Sciences, Amman, Jordan. Electronic address: m.azzeh@psut.edu.jo.
Sami AkbulutDepartment of Biostatistics and Medical Informatics, Faculty of Medicine, Inonu University, 44280, Malatya, Turkey; Inonu University, Faculty of Medicine, Department of Surgery, 44280, Malatya, Turkey; Inonu University, Faculty of Medicine, Department of Public Health, 44280, Malatya, Turkey. Electronic address: akbulutsami@gmail.com.
Inonu University · TRKing Hussein Cancer Center · JO

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

aimCOVID-19 has revealed the need for fast and reliable methods to assist clinicians in diagnosing the disease. This article presents a model that applies explainable artificial intelligence (XAI) methods based on machine learning techniques on COVID-19 metagenomic next-generation sequencing (mNGS) samples.

methodsIn the data set used in the study, there are 15,979 gene expressions of 234 patients with COVID-19 negative 141 (60.3%) and COVID-19 positive 93 (39.7%). The least absolute shrinkage and selection operator (LASSO) method was applied to select genes associated with COVID-19. Support Vector Machine - Synthetic Minority Oversampling Technique (SVM-SMOTE) method was used to handle the class imbalance problem. Logistics regression (LR), SVM, random forest (RF), and extreme gradient boosting (XGBoost) methods were constructed to predict COVID-19. An explainable approach based on local interpretable model-agnostic explanations (LIME) and SHAPley Additive exPlanations (SHAP) methods was applied to determine COVID-19- associated biomarker candidate genes and improve the final model's interpretability.

resultsFor the diagnosis of COVID-19, the XGBoost (accuracy: 0.930) model outperformed the RF (accuracy: 0.912), SVM (accuracy: 0.877), and LR (accuracy: 0.912) models. As a result of the SHAP, the three most important genes associated with COVID-19 were IFI27, LGR6, and FAM83A. The results of LIME showed that especially the high level of IFI27 gene expression contributed to increasing the probability of positive class.

conclusionsThe proposed model (XGBoost) was able to predict COVID-19 successfully. The results show that machine learning combined with LIME and SHAP can explain the biomarker prediction for COVID-19 and provide clinicians with an intuitive understanding and interpretability of the impact of risk factors in the model.

Indexed as

Artificial IntelligenceCOVID-19Calcium CompoundsGenetic MarkersHumansNeoplasm ProteinsOxidesRisk FactorsCalcium CompoundsFAM83A protein, humanGenetic MarkerslimeNeoplasm ProteinsOxidesCOVID-19Explainable artificial intelligenceLIMESHAPXGBoost

Identifiers

PMID36738712
PMCPMC9889119
OpenAlexW4318677254

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.