Evidence map›Paper›PMID 42088116›Full record

ArticleFrontiers in bioinformatics2026

Feature representation for explainable CRISPR off-target prediction and base editing efficiency.

Faiza Hasin, Michele Minervini, Corrado Mencar, Giuseppe Ventrella, Arianna Consiglio, Alessandro Orro, Tommaso Selmi

Abstract read
In one paragraph

Article in Frontiers in bioinformatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Faiza HasinDepartment of Computer Science, University of Bari Aldo Moro, Bari, Italy.
Michele MinerviniDepartment of Computer Science, University of Bari Aldo Moro, Bari, Italy.
Corrado MencarDepartment of Computer Science, University of Bari Aldo Moro, Bari, Italy.
Giuseppe VentrellaDepartment of Computer Science, University of Bari Aldo Moro, Bari, Italy.
Arianna ConsiglioInstitute for Biomedical Technologies (Bari Unit), National Research Council (CNR), Bari, Italy.
Alessandro OrroInstitute for Biomedical Technologies (Milan Unit), National Research Council (CNR), Milan, Italy.
Tommaso SelmiInstitute for Biomedical Technologies (Milan Unit), National Research Council (CNR), Milan, Italy.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Introduction: The interaction between guide RNAs (gRNAs) and target DNA sequences is a critical factor in the effectiveness of CRISPR/Cas9 (Clustered Regularly Interspaced Short Palindromic Repeats/CRISPR-associated protein 9) gene editing. Predicting these interactions accurately necessitates models that offer biological knowledge in addition to high accuracy. This study analyzes the impact of feature representation on accuracy and interpretability in off-target prediction. Methods: We address two CRISPR applications: gene knockout (KO) and base editing (BE) using distinct benchmark datasets. For the KO problem, we utilized CHANGE-seq and GUIDE-seq to evaluate paired sequence representations, while the Hanna screening dataset has been used for BE. We approached the prediction problem both as a classification and regression task using XGBoost models. Results: In the case of KO, there is not a single universally optimal encoding. For both classification and regression, One-Hot and its variants (OH, OH5C) achieve the best results on GUIDE-seq (AUPR = 0.661, Pearson = 0.756), while the Bulges representation performs best on CHANGE-seq (AUPR = 0.612, Pearson = 0.602). In the case of BE, One-hot encoding consistently outperforms K-mer representation for predictive accuracy both as regression and classification (AUPR = 0.723, Pearson = 0.746). Discussion: Our analysis demonstrates comparable predictive performance across both gene knockout and base editing tasks, confirming the robustness of the framework in distinct editing domains. Interpretability analysis using SHapley Additive exPlanations (SHAP) reveals that despite different mechanisms, the Protospacer Adjacent Motif (PAM)-proximal region remains a critical feature for prediction for both editing mechanisms.

Indexed as

base editingCRISPR/Cas9explainabilityfeature representationoff-target predictionshapXGBoost

Identifiers

PMID42088116
PMCPMC13136712

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.