Evidence map›Paper›PMID 41680846›Full record

ArticleJournal of cheminformatics2026

Integrating artificial intelligence and manual curation to enhance bioassay annotations in ChEMBL.

Ines Smit, Melissa F Adasme, Emma Manners, Sybilla Corbett, Nicolas Bosc, Hoang-My-Anh Do, Andrew R Leach, Noel M O'Boyle, Barbara Zdrazil

Erratum issuedAbstract read
In one paragraph

Article in Journal of cheminformatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. An erratum has been issued. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. AI semantics for biomedical data integration.bioRxiv : the preprint server for biology · 2026
    Article
  2. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

9 authors.

Ines Smit *European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridgeshire, CB101SD, UK.ORCID http://orcid.org/0000-0002-1772-6487
Melissa F Adasme *European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridgeshire, CB101SD, UK.ORCID http://orcid.org/0000-0003-2217-4629
Emma Manners *European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridgeshire, CB101SD, UK.ORCID http://orcid.org/0000-0001-7875-1259
Sybilla CorbettEuropean Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridgeshire, CB101SD, UK.ORCID http://orcid.org/0000-0002-1242-1481
Nicolas BoscEuropean Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridgeshire, CB101SD, UK.ORCID http://orcid.org/0000-0003-3562-1328
Hoang-My-Anh DoEuropean Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridgeshire, CB101SD, UK.
Andrew R LeachEuropean Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridgeshire, CB101SD, UK.ORCID http://orcid.org/0000-0001-8178-0253
Noel M O'BoyleEuropean Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridgeshire, CB101SD, UK.ORCID http://orcid.org/0000-0003-4879-2003
Barbara ZdrazilEuropean Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridgeshire, CB101SD, UK. bzdrazil@ebi.ac.uk.ORCID http://orcid.org/0000-0001-9395-1515

Funding

Innovative Medicines Initiative 875510Wellcome TrustWellcome Trust 104104/A/14/ZWellcome Trust 228142/Z/23/Z
6 · The paper itself

Abstract

As the volume and diversity of bioactivity data in ChEMBL continues to grow, ensuring that assay metadata is standardized, interoperable, and machine-readable is critical for effective use in cheminformatics and ML applications. In this work, we present recent efforts to enhance the quality and granularity of bioassay annotations in ChEMBL through a combination of manual and semi-manual curation and AI-driven approaches. We introduce a "perfect assay description" template to guide consistent annotation and demonstrate how natural language processing techniques and multi-class classification can be used to automatically extract key assay parameters and assign broad assay categories for legacy data. We report on the development, validation, and application of a spaCy-based NER model that identifies experimental methods with high precision and recall, as well as a complementary classification model that refines ASSAY_TYPE categorization beyond the existing schema. In addition, we describe improvements to metadata extraction for ADME endpoints, organism and protein variant annotations, and ontology linking using tools such as text2term. Together, these enhancements significantly advance the FAIRness of ChEMBL's bioassay data, enabling more robust downstream analyses and more precise compound-target activity modeling.

Indexed as

Bioassay annotationBioassay OntologyChEMBLExperimental methodFAIR dataMachine learningNamed entity recognitionNatural language processing

Identifiers

PMID41680846
PMCPMC12903245

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.