Evidence map›Paper›PMID 41502004›Full record

ArticleBioinformatics (Oxford, England)2026

A hybrid unsupervised methodology on artificial intelligence filtering for automatically processing cellular DNA-encoded library (DEL) datasets.

Yiran Huang, Xiao Tan, Xiaoyu Li, Feng Xiong, Siu Ming Yiu

Abstract read
In one paragraph

Article in Bioinformatics (Oxford, England), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Yiran HuangSchool of Pharmacy, Shenzhen University Medical School, Shenzhen University, Shenzhen 518060, China.ORCID 0009-0002-2215-3422
Xiao TanDepartment of Computer Science, The University of Hong Kong, Hong Kong SAR, 999077, China.ORCID 0000-0003-0191-9185
Xiaoyu LiDepartment of Chemistry and State Key Laboratory of Synthetic Chemistry, The University of Hong Kong, Hong Kong SAR, 999077, China.ORCID 0000-0002-8907-6727
Feng XiongDepartment of Chemistry and State Key Laboratory of Synthetic Chemistry, The University of Hong Kong, Hong Kong SAR, 999077, China.ORCID 0000-0001-7159-9522
Siu Ming YiuDepartment of Computer Science, The University of Hong Kong, Hong Kong SAR, 999077, China.ORCID 0000-0002-3975-8500

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

motivationDNA-encoded library (DEL) technology has been developed as a powerful platform for drug development. Live cell-based selection methodologies were recently developed to expedite drug candidate discovery with higher biological relevance. Nevertheless, hit characterization is challenged by prominent background signals of cell-based selections. Therefore, automated data processing streamline compatible with noisy sequencing output is highly desirable.

resultsHerein, we report an innovative automatic method that enables the most promising hit identification from large quantities of cell-based DEL datasets with improved accuracy and efficiency. This processing workflow is based on a comprehensive unsupervised algorithm incorporating data pre-processing, feature extracting and outlier filtering, descriptor-based classification, similarity score ranking, and active compound prediction. We performed methodology development with two DEL selection datasets targeting insulin receptor (INSR) on live cells, from both ∼30 million- and 1.033 billion-membered libraries. The automated scheme has demonstrated high consistency with experimental results as well as self-adaptivity to on-cell DEL datasets with varied library scales. Extended methodology application to cellular thrombopoietin receptor (TPOR) further substantiated the algorithmic generalization capability regarding target proteins. Thus, this approach can serve as a widely applicable workflow automatically differentiating hit compounds and thereby facilitates drug development from candidate discovery. AVAILABILITY AND IMPLEMENTATION: The complete datasets, source code, and pre-trained models are made available at https://doi.org/10.5281/zenodo.17452392 and https://doi.org/10.5281/zenodo.17569557.

Indexed as

Artificial IntelligenceDNADrug DiscoveryGene LibraryUnsupervised Machine LearningAlgorithmsHumansDNA

Identifiers

PMID41502004
PMCPMC12836421

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.