Evidence map›Paper›PMID 41547822›Full record

ArticleBMC medical informatics and decision making2026

Evaluation of large language models in nutrition risk screening: a comparative analysis across 8 LLMs based on real-world EHR datasets.

Si-Yu Gu, Die Yao, Yao Yao, Xing-Xing Cen, Jun-Yi Yuan

Abstract readComparative Study
In one paragraph

Article in BMC medical informatics and decision making, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Si-Yu Gu *Shanghai Chest Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China.
Die Yao *Shanghai Chest Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China.
Yao YaoShanghai Chest Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China.
Xing-Xing CenShanghai Chest Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China. cxx2347@163.com.
Jun-Yi YuanShanghai Chest Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China. yuanjunyi_yjy@163.com.

Funding

Shanghai Chest Hospital XKRGZN202509Shanghai Special Project for Promoting High-quality Industrial Development 2024-GZL-RGZN-02011
6 · The paper itself

Abstract

backgroundNutrition risk screening (NRS) is a critical step in the early identification of malnutrition among hospitalized patients. Traditional methods, which rely on manual assessments using tools such as Nutrition Risk Screening 2002 (NRS-2002) based on electronic health records (EHRs), are time-consuming and often yield in accuracy. Large language models (LLMs) offer promising potential to automate this process; however, their capabilities in this scenario remain underexplored and not yet fully realized.

methodsA multidisciplinary expert group developed standardized scoring criteria and structured prompts using prompt engineering techniques, optimized using an 80-case prompt development cohort. Eight advanced LLMs with different architectures, parameter scales, and openness levels were evaluated using 592 real-world inpatient EHRs. Each model independently assessed every case twice with uniform structured prompt, determined nutritional risk, and generated reasoning outputs decomposed into total and domain-level scores for nutritional status, disease severity, and age. Model performance was assessed across multiple dimensions, including accuracy: risk-specific, total-specific, and domain-specific correct assessment rate (CrAR), consistency: consistent assessment rate (CsAR), and efficiency: processing time.

resultsUsing a structured prompt, five of eight LLMs achieved over 90% CrAR in binary nutritional risk classification, with top models reaching 99.16% (DeepSeek-R1-671B). Performance in total score CrAR varied widely (54.73% − 95.60%), while domain-specific CrAR was highest in nutritional status, with serum albumin and age scoring near-perfect across models. The CrAR of disease severity was more challenging, showing greater inter-model variability. Larger parameter scales LLMs demonstrated higher accuracy and repeatability, with Cohen’s κ up to 0.99, whereas smaller LLMs like Qwen3-8B showed marked declines (κ = 0.69). Domain-level consistency was particularly strong in structured subdomains (albumin and age), while subdomains requiring complex clinical inference (disease burden) yielded lower consistency. Qwen3-235B-A22B-Thinking-2507 was lowest (60.7s/case); smaller LLMs had lower accuracy and faster responses.

conclusionsLLMs guided by structured prompts can effectively perform automated NRS, with larger parameter scales models achieving near-expert reliability. These findings support the integration of LLMs into clinical workflows, especially in settings with limited human resources. Future work should explore fine-tuning smaller LLMs for greater deployment efficiency while maintaining diagnostic robustness, as well as expanding applications of LLMs to broader clinical decision-making tasks in support of health equity.

Indexed as

Electronic Health RecordsLarge Language ModelsMalnutritionNutrition AssessmentHumansRisk AssessmentArtificial intelligenceElectronic health recordsHealth equityLarge language modelsNutrition risk screening

Identifiers

PMID41547822
PMCPMC12896154

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.