Evidence map›Paper›PMID 41251541›Full record

ArticleJournal of medical Internet research2025

Automated Multitier Tagging of Chinese Online Health Education Resources Using a Large Language Model: Development and Validation Study.

Jialin Meng, Ruiming Dai, Xiaolan Huang, Yi Gu, Shixing Yan, Xiaoke Wang, Jingrong Gao, Tian-Tian Zhang

Abstract readValidation Study
In one paragraph

Article in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Jialin MengSchool of Public Health, Fudan University, Shanghai, China.ORCID https://orcid.org/0009-0009-9407-6885
Ruiming DaiShanghai Center for Emerging Technologies Governance in Medicine and Public Health, Shanghai, China.ORCID https://orcid.org/0009-0007-2897-6629
Xiaolan HuangShanghai Municipal Center for Health Promotion, Shanghai, China.ORCID https://orcid.org/0009-0003-8550-9761
Yi GuShanghai Municipal Center for Health Promotion, Shanghai, China.ORCID https://orcid.org/0000-0003-0138-7502
Shixing YanShanghai Center for Emerging Technologies Governance in Medicine and Public Health, Shanghai, China.ORCID https://orcid.org/0009-0009-7674-1853
Xiaoke WangShanghai Center for Emerging Technologies Governance in Medicine and Public Health, Shanghai, China.ORCID https://orcid.org/0009-0008-7758-2308
Jingrong GaoShanghai Municipal Center for Health Promotion, Shanghai, China.ORCID https://orcid.org/0009-0001-6615-6687
Tian-Tian ZhangSchool of Public Health, Fudan University, Shanghai, China.ORCID https://orcid.org/0000-0003-1730-4185

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundPrecision health promotion, which aims to tailor health messages to individual needs, is hampered by the lack of structured metadata in vast digital health resource libraries. This bottleneck prevents scalable, personalized content delivery and exacerbates information overload for the public.

objectiveThis study aimed to develop, deploy, and validate an automated tagging system using a large language model (LLM) to create the foundational metadata infrastructure required for tailored health communication at scale.

methodsWe developed a comprehensive, 3-tier health promotion taxonomy (10 primary, 34 secondary, and 90,562 tertiary tags) using a hybrid Delphi and corpus-mining methodology. We then constructed a hybrid inference pipeline by fine-tuning a Baichuan2-7B LLM with low-rank adaptation for initial tag generation. This was then refined by a domain-specific named entity recognition model and standardized against a vector database. The system's performance was evaluated against manual annotations from nonexpert staff on a test set of 1000 resources. We used a "no gold standard" framework, comparing the artificial intelligence-human (A-H) interrater reliability (IRR) with a supplemental human-human (H-H) IRR baseline and expert adjudication for cases where artificial intelligence provided additional tags ("AI Additive").

resultsThe A-H agreement was moderate (Cohen κ=0.54, 95% CI 0.53-0.56; Jaccard similarity coefficient=0.48, 95% CI 0.46-0.50). Critically, this was higher than the baseline nonexpert H-H agreement (Cohen κ=0.32, 95% CI 0.29-0.35; Jaccard similarity coefficient=0.35, 95% CI 0.27-0.43). A granular analysis of disagreements revealed that in 15.9% (159/1000) of the cases, the "AI Additive" tags were not identified by human annotators. Expert adjudication of these cases confirmed that the "AI Additive" tags were correct and relevant with a precision of 90% (45/50; 95% CI 78.2%-96.7%).

conclusionsA fine-tuned LLM, integrated into a hybrid pipeline, can function as a powerful augmentation tool for health content annotation. The system's consistency (A-H κ=0.54) was found to be superior to the baseline human workflow (H-H κ=0.32). By moving beyond simple automation to reliably identify relevant health topics missed by manual annotators with high, expert-validated accuracy, this study provides a robust technical and methodological blueprint for implementing artificial intelligence to enhance precision health communication in public health settings.

Indexed as

Health EducationInternetLanguageArtificial IntelligenceChinaHumansLarge Language ModelsChinadigital healthhealth promotionlarge language modelnamed entity recognitionnatural language processingtagging

Identifiers

PMID41251541
PMCPMC12756663

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.