Evidence map›Paper›PMID 41043140›Full record

ArticleJMIR formative research2025

Identification of Syndrome Types in Patients With Pancreatic Cancer From Free Text in Electronic Medical Records: Model Development and Validation.

He Ba, Haojie Du, Chienshan Cheng, Yuan Zhang, Linjie Ruan, Zhen Chen

Abstract readValidation Study
In one paragraph

Article in JMIR formative research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Trial
  2. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

He Ba *Qingdao Institute, Department of Integrative Oncology, Fudan University Shanghai Cancer Center, Qingdao, China.ORCID 0000-0003-1572-3108
Haojie Du *Key Laboratory of Surface & Interface Science of Polymer Materials of Zhejiang Province, School of Chemistry and Chemical Engineering, Zhejiang Sci-Tech University, Hangzhou, China.ORCID 0009-0006-1313-5184
Chienshan ChengDepartment of Oncology, Shanghai Medical College, Fudan University, Shanghai, China.ORCID 0000-0003-4885-6759
Yuan ZhangDepartment of Oncology, Shanghai Medical College, Fudan University, Shanghai, China.ORCID 0000-0001-7497-2513
Linjie RuanDepartment of Oncology, Shanghai Medical College, Fudan University, Shanghai, China.ORCID 0009-0003-1220-2371
Zhen ChenDepartment of Oncology, Shanghai Medical College, Fudan University, Shanghai, China.ORCID 0000-0002-4502-0801

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundSyndrome differentiation is crucial in traditional Chinese medicine (TCM) diagnosis and treatment, but it heavily relies on expert experience, limiting systematic standardization.

objectiveThis study developed and validated a BERT (bidirectional encoder representations from transformers)-based model, the traditional Chinese medicine pancreatic cancer syndrome differentiation bidirectional encoder representations from transformers (TCMPCSD-BERT), using in-house pancreatic cancer medical records, to digitalize expert knowledge and support standardized syndrome differentiation in TCM.

methodsA retrospective dataset of pancreatic cancer cases (2011-2024) from Fudan University Shanghai Cancer Center was annotated into 4 TCM syndrome types by 2 experts (Cohen κ=0.913). The proposed TCMPCSD-BERT model was compared with conventional models (long short-term memory and text convolutional neural network) embedded in TCM diagnostic tools and with large language models (LLMs; ChatGPT-4o, Kimi, Ernie Bot 4.0 Turbo, and Zhipu Qingyan) under a prompt engineering framework. Performance evaluation on in-house data was supplemented with attention visualizations and integrated gradients analyses for interpretability. The McNemar test assessed classification accuracy differences, while bootstrap 95% CIs quantified statistical uncertainty and stability. The Welch t test (2-tailed) was used to evaluate mean differences between TCMPCSD-BERT and the comparator models.

resultsAmong 6830 records, case counts were damp-heat syndrome (n=1694), spleen-deficiency syndrome (n=1185), damp-heat with spleen-deficiency syndrome (n=1178), and others (n=2773). On the test set, McNemar test showed significantly higher accuracy for TCMPCSD-BERT than the 3 baseline models and generally better performance than LLMs. In all comparisons, TCMPCSD-BERT achieved higher mean macroprecision, macrorecall, macro-F

conclusionsThe TCMPCSD-BERT model shows potential for automated syndrome differentiation from unstructured clinical texts and broader application in TCM. Based on this study, it appears to improve operability over 4-diagnostic instruments and web-based platforms, and offers greater stability and accuracy than LLMs in specific tasks. However, these findings should be interpreted cautiously, given the subjectivity of syndrome definitions, data imbalance, and reliance on preprocessed, expert-annotated data. Further studies involving larger and more diverse populations are needed to validate its generalizability and support its broader application in real-world settings.

Indexed as

Electronic Health RecordsMedicine, Chinese TraditionalPancreatic NeoplasmsDiagnosis, DifferentialFemaleHumansMaleNeural Networks, ComputerRetrospective StudiesSyndromeBERTbidirectional encoder representations from transformersnatural language processingpancreatic cancersyndrome differentiationtraditional Chinese medicine

Identifiers

PMID41043140
PMCPMC12534766

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.