Evidence map›Paper›PMID 39292535›Full record

ArticleBioinformatics (Oxford, England)2024

AutoPeptideML: a study on how to build more trustworthy peptide bioactivity predictors.

Raúl Fernández-Díaz, Rodrigo Cossio-Pérez, Clement Agoni, Hoang Thanh Lam, Vanessa Lopez, Denis C Shields

Abstract read
In one paragraph

Article in Bioinformatics (Oxford, England), 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 11 papers.

0numbers the graph read from it
0cells of the map it votes in
11citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

11 citing papers in PubMed.

  1. Review
  2. Article
  3. Review
  4. Article
  5. Article
  6. Article
  7. Review
  8. Review
  9. Article
  10. Article
  11. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Raúl Fernández-DíazIBM Research, Dublin, Dublin D15 HN66, Ireland.
Rodrigo Cossio-PérezSchool of Medicine, University College Dublin, Dublin D04 C1P1, Ireland.
Clement AgoniSchool of Medicine, University College Dublin, Dublin D04 C1P1, Ireland.
Hoang Thanh LamIBM Research, Dublin, Dublin D15 HN66, Ireland.
Vanessa LopezIBM Research, Dublin, Dublin D15 HN66, Ireland.
Denis C ShieldsSchool of Medicine, University College Dublin, Dublin D04 C1P1, Ireland.ORCID 0000-0003-4015-2474

Funding

European Union's Horizon 2020 research and innovation programmeScience Foundation Ireland
6 · The paper itself

Abstract

motivationAutomated machine learning (AutoML) solutions can bridge the gap between new computational advances and their real-world applications by enabling experimental scientists to build their own custom models. We examine different steps in the development life-cycle of peptide bioactivity binary predictors and identify key steps where automation cannot only result in a more accessible method, but also more robust and interpretable evaluation leading to more trustworthy models.

resultsWe present a new automated method for drawing negative peptides that achieves better balance between specificity and generalization than current alternatives. We study the effect of homology-based partitioning for generating the training and testing data subsets and demonstrate that model performance is overestimated when no such homology correction is used, which indicates that prior studies may have overestimated their performance when applied to new peptide sequences. We also conduct a systematic analysis of different protein language models as peptide representation methods and find that they can serve as better descriptors than a naive alternative, but that there is no significant difference across models with different sizes or algorithms. Finally, we demonstrate that an ensemble of optimized traditional machine learning algorithms can compete with more complex neural network models, while being more computationally efficient. We integrate these findings into AutoPeptideML, an easy-to-use AutoML tool to allow researchers without a computational background to build new predictive models for peptide bioactivity in a matter of minutes. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, and data are available at https://github.com/IBM/AutoPeptideML and a dedicated web-server at http://peptide.ucd.ie/AutoPeptideML. A static version of the software to ensure the reproduction of the results is available at https://zenodo.org/records/13363975.

Indexed as

AlgorithmsMachine LearningPeptidesComputational BiologyDatabases, ProteinNeural Networks, ComputerSoftwarePeptides

Identifiers

PMID39292535
PMCPMC11438549

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.