Evidence map›Paper›PMID 40210962›Full record

ArticleScientific reports2025

Comparing large scale and selected feature learning for community acquired pneumonia prognosis prediction using clinical data: a stacked ensemble approach.

Ji Hyun Lee, Hyun Woo Lee, Hyo Jin Lee, Tae Yun Park, Kwang Nam Jin, Dong Hyun Kim, Borim Ryu

Abstract readComparative Study
In one paragraph

Article in Scientific reports, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Trial
  2. Predicting macrolide resistance in pediatric Mycoplasma pneumoniae pneumonia: A machine learning modeling study.European journal of clinical microbiology & infectious diseases : official publication of the European Society of Clinical Microbiology · 2026
    Article
  3. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Ji Hyun Lee *Department of Radiology, Seoul Metropolitan Government-Seoul National University Boramae Medical Center, 20, Boramae-ro 5-gil, Dongjak-gu, Seoul, Republic of Korea.
Hyun Woo Lee *Division of Pulmonary and Critical Care Medicine, Department of Internal Medicine, Seoul National University College of Medicine, Seoul Metropolitan Government-Seoul National University Boramae Medical Center, 20, Boramae-ro 5-gil, Dongjak-gu, Seoul, Republic of Korea.
Hyo Jin LeeDivision of Pulmonary and Critical Care Medicine, Department of Internal Medicine, Seoul National University College of Medicine, Seoul Metropolitan Government-Seoul National University Boramae Medical Center, 20, Boramae-ro 5-gil, Dongjak-gu, Seoul, Republic of Korea.
Tae Yun ParkDivision of Pulmonary and Critical Care Medicine, Department of Internal Medicine, Seoul National University College of Medicine, Seoul Metropolitan Government-Seoul National University Boramae Medical Center, 20, Boramae-ro 5-gil, Dongjak-gu, Seoul, Republic of Korea.
Kwang Nam JinDepartment of Radiology, Seoul Metropolitan Government-Seoul National University Boramae Medical Center, 20, Boramae-ro 5-gil, Dongjak-gu, Seoul, Republic of Korea.
Dong Hyun KimDepartment of Radiology, Seoul Metropolitan Government-Seoul National University Boramae Medical Center, 20, Boramae-ro 5-gil, Dongjak-gu, Seoul, Republic of Korea. mi4ri4@gmail.com.
Borim RyuCenter for Data Science, Biomedical Research Institute Seoul Metropolitan Government-Seoul National University Boramae Medical Center, 20, Boramae-ro 5-gil, Dongjak-gu, Seoul, Republic of Korea. borim.ryu@gmail.com.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

This study investigated and validated all-cause in-hospital death prediction models for hospitalized pneumonia patients based on large-scale clinical data, including diagnoses, medication prescriptions, and laboratory test codes. Feature selection was performed using both large-scale feature learning with a Common Data Model (CDM) and specific pneumonia-related risk factors. A stacked ensemble mixed machine-learning model was compared with traditional machine-learning models. Accuracy, F1-score, the Area Under Precision Recall Curve (AUPRC) and the Area Under the Receiver Operating Characteristic (AUROC) were used for performance evaluation. For large-scale feature learning using a CDM, the ensemble model (LASSO LR + GBM + RF) achieved the highest performance. For the 365-day lookback, the ensemble model's AUROC was 0.867 (95% CI: 0.823-0.910), and for the 7-day lookback (AUROC 0.867, 95% CI: 0.822-0.912). In contrast, for feature learning based on selected pneumonia risk factors, among the traditional models, the RF model performed best with AUROCs of 0.774 (95% CI: 0.717-0.830) for the 365-day lookback and 0.773 (95% CI: 0.717-0.828) for the 7-days lookback. Leveraging large-scale feature learning within the CDM and using a stacked ensemble model predicts more accurately and robustly, highlighting the potential to capture complex relationships among clinical features and improve prognostic assessments.

Indexed as

Community-Acquired InfectionsMachine LearningPneumoniaAgedAged, 80 and overCommunity-Acquired PneumoniaFemaleHospital MortalityHumansMaleMiddle AgedPrognosisRisk FactorsROC CurveCommon data modelMortality predictionPneumoniaStacked ensemble modelTreatment planning

Identifiers

PMID40210962
PMCPMC11985930

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.