Evidence map›Paper›PMID 42289972›Full record

ArticleBioinformatics (Oxford, England)2026

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

Patryk Jarnot, Joanna Ziemska-Legiecka, Marcin Grynberg, Vasilis J Promponas, Aleksandra Gruca

Abstract read
In one paragraph

Article in Bioinformatics (Oxford, England), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Patryk JarnotDepartment of Computer Networks and Systems, Silesian University of Technology, Gliwice, 44-100, Poland.ORCID 0000-0002-8318-2270
Joanna Ziemska-LegieckaInstitute of Biochemistry and Biophysics, Polish Academy of Sciences, Warsaw, 02-106, Poland.
Marcin GrynbergInstitute of Biochemistry and Biophysics, Polish Academy of Sciences, Warsaw, 02-106, Poland.
Vasilis J PromponasBioinformatics Research Laboratory, Department of Biological Sciences, University of Cyprus, 1 University Avenue, Nicosia, 2109, Cyprus.ORCID 0000-0003-3352-4831
Aleksandra GrucaDepartment of Computer Networks and Systems, Silesian University of Technology, Gliwice, 44-100, Poland.ORCID 0000-0003-2337-1894

Funding

COST Action CA21160National Science Centre 2020/39/B/ST6/03447
6 · The paper itself

Abstract

motivationShort tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools.

resultsHere, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Indexed as

Computational BiologyMicrosatellite RepeatsProteinsSequence Analysis, ProteinAlgorithmsAmino Acid SequenceCluster AnalysisClustering AlgorithmsSoftwareProteins

Identifiers

PMID42289972
PMCPMC13360283

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.