Evidence map›Paper›PMID 42113882›Full record

ArticlePLoS computational biology2026

scHilda: Hierarchical Integration of LLM with KG database for single cell type annotation.

Yilang Li, Yidi Sun, Aoyun Geng, Junlin Xu, Yajie Meng, Feifei Cui, Leyi Wei, Quan Zou, Xiulai Li, Zilong Zhang

Abstract read
In one paragraph

Article in PLoS computational biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Yilang LiSchool of Cyberspace Security (School of Cryptology), Hainan University, Haikou, China.
Yidi SunSchool of Computer Science and Technology, Hainan University, Haikou, China.
Aoyun GengSchool of Computer Science and Technology, Hainan University, Haikou, China.
Junlin XuSchool of Computer Science and Technology, Wuhan University of Science and Technology, Wuhan, Hubei, China.
Yajie MengSchool of Computer Science and Artificial Intelligence, Wuhan Textile University, Wuhan, Hubei, China.
Feifei CuiSchool of Computer Science and Technology, Hainan University, Haikou, China.
Leyi WeiCentre for Artificial Intelligence driven Drug Discovery, Faculty of Applied Science, Macao Polytechnic University, Macao Special Administrative Region, China.
Quan ZouInstitute of Fundamental and Frontier Sciences, University of Electronic Science and Technology of China, Chengdu, China.ORCID https://orcid.org/0000-0001-6406-1142
Xiulai LiSchool of Computer Science and Technology, Hainan University, Haikou, China.
Zilong ZhangSchool of Computer Science and Technology, Hainan University, Haikou, China.ORCID https://orcid.org/0000-0002-4934-1258

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Cell type annotation in single-cell RNA sequencing is a critical bottleneck, with existing automated methods facing limitations in accuracy, interpretability, and generalization to novel cell types. Although Large Language Models (LLMs) have recently shown potential in single-cell annotation, they are prone to inherent "hallucinations". Furthermore, a critical challenge is utilizing imperfect and potentially noisy external knowledge bases in a principled and robust manner to effectively constrain and enhance the LLM's reasoning capabilities. To address this, we propose scHilda, a novel framework designed to tackle this challenge. It deeply integrates an external Knowledge Graph into the LLM's reasoning process and employs a hierarchical arbitration annotation strategy. This strategy first identifies major cell lineages with the support of a global knowledge base and then dynamically retrieves focused subgraph domain information related to that lineage to precisely resolve cell subtypes. This dynamic knowledge-enhanced reasoning mechanism effectively constrains the LLM's decision space, reduces the risk of hallucination, and mitigates potential misguidance from knowledge base deficiencies. Tests on multiple benchmark datasets show that scHilda outperforms existing methods, achieving state-of-the-art (SOTA) performance. Notably, scHilda demonstrates exceptional robustness when handling complex mixed samples and enables lower-cost lightweight LLMs to achieve annotation performance close to that of top-tier models. Furthermore, rigorous statistical evaluations, alongside detailed interpretability case studies and query complexity analyses, validate the framework's efficiency and transparent decision-making. By deeply integrating the reasoning power of LLMs with structured biological knowledge, scHilda not only improves the accuracy and interpretability of cell annotation but also provides a new paradigm for building the next generation of trustworthy biological AI systems.

Indexed as

Computational BiologyMolecular Sequence AnnotationSingle-Cell AnalysisAlgorithmsAnimalsCell LineageHumansKnowledge BasesLarge Language ModelsSoftware

Identifiers

PMID42113882
PMCPMC13175463

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.