Evidence map›Paper›PMID 40203463›Full record

ArticleInternational journal of medical informatics2025

Predicting hospital admissions, ICU utilization, and prolonged length of stay among febrile pediatric emergency department patients using incomplete and imbalanced electronic health record (EHR) data strategies.

Tom Velez, Zara Ibrahim, Kanayo Duru, Dante Velez, Maria Triantafyllou, Kenneth McKinley, Pasha Saif, Panagiotis Kratimenos, Andy Clark, Ioannis Koutroulis

Abstract read
In one paragraph

Article in International journal of medical informatics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. An explainable machine learning approach to predicting carbapenem resistance inFrontiers in cellular and infection microbiology · 2026
    Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Tom VelezComputer Technology Associates, Cardiff, CA, United States.
Zara IbrahimDepartment of Pediatrics, Children's National Hospital, Washington, DC, United States.
Kanayo DuruDepartment of Pediatrics, Children's National Hospital, Washington, DC, United States; Brown University, Providence, RI, United States.
Dante VelezDepartment of Pediatrics, Children's National Hospital, Washington, DC, United States.
Maria TriantafyllouCenter for Genetic Medicine Research, Children's National Research Institute, Washington, DC, United States.
Kenneth McKinleyDepartment of Pediatrics, Children's National Hospital, Washington, DC, United States; George Washington University School of Medicine and Health Sciences, Washington, DC, United States.
Pasha SaifVirginia Tech Carilion School of Medicine, Roanoke, VA, United States.
Panagiotis KratimenosDepartment of Pediatrics, Children's National Hospital, Washington, DC, United States; George Washington University School of Medicine and Health Sciences, Washington, DC, United States.
Andy ClarkComputer Technology Associates, Cardiff, CA, United States.
Ioannis KoutroulisDepartment of Pediatrics, Children's National Hospital, Washington, DC, United States; Center for Genetic Medicine Research, Children's National Research Institute, Washington, DC, United States; George Washington University School of Medicine and Health Sciences, Washington, DC, United States. Electronic address: ikoutrouli@childrensnational.org.

Funding

Biomarker-enhanced Artificial Intelligence-Based Pediatric Sepsis Screening Tool Towards Early Recognition and Personalized TherapeuticsR41AI167224 · NIAID · COMPUTER TECHNOLOGY ASSOCIATES, INC. · PI KOUTROULIS, IOANNIS, VELEZ, CARMELO ELLIOT · 2022 to 2023
$597k
NIAID NIH HHS R41 AI167224
6 · The paper itself

Abstract

objectiveDetermine the efficacy of commonly used approaches to handling missing and/or imbalanced Electronic Health Record (EHR) data on the performance of predictive models targeting risk of admission, intensive care unit (ICU) use, or prolonged length of stay (PLOS) among presenting febrile pediatric emergency department (ED) patients. MATERIALS AND

methodsHistorical ED EHR data was used to train a series of XGBoost (XGB) and logistic regression (LR) classifiers. Data handling strategies included imputation methods (multiple imputation (MI), median imputation, complete case (CC) analysis), and imbalanced data corrections (minority oversampling, stratified sub-group analysis). Model performance was evaluated using discriminative (AUC, AUPRC) and calibration metrics (Brier score, Z-scores, p-values).

resultsAmong the study population, 34 % were admitted, 2 % utilized the ICU, and 7 % had a PLOS. Significant data missingness was observed and determined to be not at random (MNAR). In predicting admissions using data recorded within the first two hours of presentation, LR trained using full cohort with median imputation was comparable to MI yielding well-calibrated admissions models with an AUC/AUPRC of 0.82/0.73 while CC analysis yielded an AUC/AUPRC of 0.76/0.78. XGB, trained with unimputed data, produced a well-calibrated admissions classifier with an AUC/AUPRC of 0.85/0.78. In contrast, imbalanced data correction techniques, including synthetic minority oversampling (SMOTE), risk stratification, or the use of XGB did not significantly improve the poor AUPRC and calibration performance of LR models predicting ICU and PLOS.

conclusionBoth XGB and LR with median imputation demonstrated robust performance in predicting admissions in the presence of missing data. However, deriving clinically useful models for rare outcomes, such as ICU use or PLOS, remains a challenge due to poor precision/recall and calibration performance. Further research is needed to improve the prediction of rare outcomes in this population.

Indexed as

Electronic Health RecordsEmergency Service, HospitalFeverHospitalizationIntensive Care UnitsLength of StayPatient AdmissionChildChild, PreschoolFemaleHumansInfantMaleFebrileImbalanced dataImputationMachine learningPediatric emergency medicine

Identifiers

PMID40203463
PMCPMC12977964

What OpenQuestion holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.