Evidence map›Paper›PMID 40628826›Full record

ArticleScientific reports2025

Advancing the accuracy of clathrin protein prediction through multi-source protein language models.

Watshara Shoombuatong, Nalini Schaduangrat, Pakpoom Mookdarsanit, Jaru Nikom, Lawankorn Mookdarsanit

Abstract read
In one paragraph

Article in Scientific reports, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Watshara ShoombuatongCenter for Research Innovation and Biomedical Informatics, Faculty of Medical Technology, Mahidol University, Bangkok, 10700, Thailand. watshara.sho@mahidol.ac.th.
Nalini SchaduangratCenter for Research Innovation and Biomedical Informatics, Faculty of Medical Technology, Mahidol University, Bangkok, 10700, Thailand.
Pakpoom MookdarsanitComputer Science and Artificial Intelligence, Faculty of Science, Chandrakasem Rajabhat University, Bangkok, 10900, Thailand.
Jaru NikomResearch Methodology and Data Analytics Program, Faculty of Science and Technology, Prince of Songkla University, Pattani, 94000, Thailand.
Lawankorn MookdarsanitBusiness Information System, Faculty of Management Science, Chandrakasem Rajabhat University, Bangkok, 10900, Thailand. lawankorn.s@chandra.ac.th.

Funding

National Research Council of Thailand and Mahidol University N42A660380
6 · The paper itself

Abstract

Clathrin is a key cytoplasmic protein that serves as the predominant structural element in the formation of coated vesicles. Specifically, clarithin enables the scission of newly formed vesicles from the plasma membrane's cytoplasmic face. Efficient and accurate identification of clathrins is essential for understanding human diseases and aiding drug target development. Recent advancements in computational methods for identifying clathrins using sequence data have greatly improved large-scale clathrin screening. Here, we propose a high-accuracy computational approach, termed PLM-CLA, to achieve more accurate identification of clathrins. In PLM-CLA, we leveraged multi-source pre-trained protein language models (PLMs), which were trained on large-scale protein sequences from multiple database sources, including ProtT5-BFD, ProtT5-UR50, ProstT5, and ESM-2. These models were used to encode complementary feature embeddings, capturing diverse and valuable information. To the best of our knowledge, PLM-CLA is the first attempt designed using various PLM-based embeddings to identify clathrins. To enhance prediction performance, we utilized a feature selection method to optimize these fused feature embeddings. Finally, we employed a long short-term memory (LSTM) neural network model coupled with the optimal feature subset to identify clathrins. Benchmarking experiments, including independent tests, showed that PLM-CLA significantly outperformed state-of-the-art methods, achieving an accuracy of 0.961, MCC of 0.917, and AUC of 0.997. Furthermore, PLM-CLA secured outstanding performance in terms of MCC, with values of 0.971 and 0.904 on two existing independent test datasets. We anticipate that the proposed PLM-CLA model will serve as a promising tool for large-scale identification of clathrins in resource-limited settings.

Indexed as

ClathrinComputational BiologyAlgorithmsDatabases, ProteinHumansNeural Networks, ComputerClathrinBioinformaticsClathrinDeep learningFeature selectionMachine learningProtein language modelSequence analysis

Identifiers

PMID40628826
PMCPMC12238356

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.