Evidence map›Paper›PMID 40705933›Full record

ArticleJMIR medical informatics2025

A Weighted Voting Approach for Traditional Chinese Medicine Formula Classification Using Large Language Models: Algorithm Development and Validation Study.

Zhe Wang, Keqian Li, Suyuan Peng, Lihong Liu, Xiaolin Yang, Keyu Yao, Heinrich Herre, Yan Zhu

Abstract readValidation Study
In one paragraph

Article in JMIR medical informatics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Zhe WangInstitute of Basic Medical Sciences, Chinese Academy of Medical Sciences; School of Basic Medicine, Peking Union Medical College, Beijing, China.ORCID 0009-0000-2387-784X
Keqian LiSchool of Medical Information, Changchun University of Chinese Medicine, Changchun, China.ORCID 0009-0002-5956-3038
Suyuan PengInstitute of Information on Traditional Chinese Medicine, China Academy of Chinese Medical Sciences, No 16, Nanxiao Street, Dongzhimen, Beijing, 100010, China, 86 010 64089639.ORCID 0000-0002-8221-7574
Lihong LiuInstitute of Information on Traditional Chinese Medicine, China Academy of Chinese Medical Sciences, No 16, Nanxiao Street, Dongzhimen, Beijing, 100010, China, 86 010 64089639.ORCID 0009-0004-8250-2772
Xiaolin YangInstitute of Basic Medical Sciences, Chinese Academy of Medical Sciences; School of Basic Medicine, Peking Union Medical College, Beijing, China.ORCID 0000-0001-9008-6650
Keyu YaoInstitute of Information on Traditional Chinese Medicine, China Academy of Chinese Medical Sciences, No 16, Nanxiao Street, Dongzhimen, Beijing, 100010, China, 86 010 64089639.ORCID 0000-0003-2655-9243
Heinrich HerreInstitute for Medical Informatics, Statistics and Epidemiology, University of Leipzig, Leipzig, Germany.ORCID 0000-0001-5343-9218
Yan ZhuInstitute of Information on Traditional Chinese Medicine, China Academy of Chinese Medical Sciences, No 16, Nanxiao Street, Dongzhimen, Beijing, 100010, China, 86 010 64089639.ORCID 0000-0002-5592-8258

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Several clinical cases and experiments have demonstrated the effectiveness of traditional Chinese medicine (TCM) formulas in treating and preventing diseases. These formulas contain critical information about their ingredients, efficacy, and indications. Classifying TCM formulas based on this information can effectively standardize TCM formulas management, support clinical and research applications, and promote the modernization and scientific use of TCM. To further advance this task, TCM formulas can be classified using various approaches, including manual classification, machine learning, and deep learning. Additionally, large language models (LLMs) are gaining prominence in the biomedical field. Integrating LLMs into TCM research could significantly enhance and accelerate the discovery of TCM knowledge by leveraging their advanced linguistic understanding and contextual reasoning capabilities. Objective: The objective of this study is to evaluate the performance of different LLMs in the TCM formula classification task. Additionally, by employing ensemble learning with multiple fine-tuned LLMs, this study aims to enhance classification accuracy. Methods: The data for the TCM formula were manually refined and cleaned. We selected 10 LLMs that support Chinese for fine-tuning. We then employed an ensemble learning approach that combined the predictions of multiple models using both hard and weighted voting, with weights determined by the average accuracy of each model. Finally, we selected the top 5 most effective models from each series of LLMs for weighted voting (top 5) and the top 3 most accurate models of 10 for weighted voting (top 3). Results: A total of 2441 TCM formulas were curated manually from multiple sources, including the Coding Rules for Chinese Medicinal Formulas and Their Codes, the Chinese National Medical Insurance Catalog for proprietary Chinese medicines, textbooks of TCM formulas, and TCM literature. The dataset was divided into a training set of 1999 TCM formulas and test set of 442 TCM formulas. The testing results showed that Qwen-14B achieved the highest accuracy of 75.32% among the single models. The accuracy rates for hard voting, weighted voting, weighted voting (top 5), and weighted voting (top 3) were 75.79%, 76.47%, 75.57%, and 77.15%, respectively. Conclusions: This study aims to explore the effectiveness of LLMs in the TCM formula classification task. To this end, we propose an ensemble learning method that integrates multiple fine-tuned LLMs through a voting mechanism. This method not only improves classification accuracy but also enhances the existing classification system for classifying the efficacy of TCM formula.

Indexed as

AlgorithmsDrugs, Chinese HerbalLanguageMedicine, Chinese TraditionalHumansLarge Language ModelsMachine LearningVotingDrugs, Chinese Herbalalgorithm developmentensemble learninglarge language modelsTCM formula classificationtraditional Chinese medicine

Identifiers

PMID40705933
PMCPMC12292024

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.