Evidence map›Paper›PMID 39146534›Full record

ArticleJMIR public health and surveillance2024

Predicting Long COVID in the National COVID Cohort Collaborative Using Super Learner: Cohort Study.

Zachary Butzin-Dozier, Yunwen Ji, Haodong Li, Jeremy Coyle, Junming Shi, Rachael V Phillips, Andrew N Mertens, Romain Pirracchio, Mark J van der Laan, Rena C Patel and 3 more

Abstract read
In one paragraph

Article in JMIR public health and surveillance, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers.

0numbers the graph read from it
0cells of the map it votes in
6citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

6 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

13 authors.

Zachary Butzin-DozierDivision of Biostatistics, University of California Berkeley School of Public Health, Berkeley, CA, United States.ORCID 0000-0001-6419-0008
Yunwen JiDivision of Biostatistics, University of California Berkeley School of Public Health, Berkeley, CA, United States.ORCID 0000-0002-6077-5102
Haodong LiDivision of Biostatistics, University of California Berkeley School of Public Health, Berkeley, CA, United States.ORCID 0000-0003-4069-5743
Jeremy CoyleDivision of Biostatistics, University of California Berkeley School of Public Health, Berkeley, CA, United States.ORCID 0000-0002-9874-6649
Junming ShiDivision of Biostatistics, University of California Berkeley School of Public Health, Berkeley, CA, United States.ORCID 0000-0002-1520-8387
Rachael V PhillipsDivision of Biostatistics, University of California Berkeley School of Public Health, Berkeley, CA, United States.ORCID 0000-0002-8474-591X
Andrew N MertensDivision of Biostatistics, University of California Berkeley School of Public Health, Berkeley, CA, United States.ORCID 0000-0002-1050-6721
Romain PirracchioDepartment of Anesthesia and Perioperative Care, University of California San Francisco, San Francisco, CA, United States.ORCID 0000-0002-7007-4908
Mark J van der LaanDivision of Biostatistics, University of California Berkeley School of Public Health, Berkeley, CA, United States.ORCID 0000-0003-1432-5511
Rena C PatelDepartment of Infectious Diseases, University of Alabama at Birmingham School of Medicine, Birmingham, AL, United States.ORCID 0000-0001-9893-5856
John M ColfordDivision of Biostatistics, University of California Berkeley School of Public Health, Berkeley, CA, United States.ORCID 0000-0002-3288-6956
Alan E HubbardDivision of Biostatistics, University of California Berkeley School of Public Health, Berkeley, CA, United States.ORCID 0000-0002-3769-0127
National COVID Cohort Collaborative (N3C) Consortium *Members are listed at the end of the manuscript, .

Funding

CD2H - The National COVID Cohort Collaborative (N3C) IDeA CTR CollaborationU24TR002306 · NCATS · UNIVERSITY OF COLORADO DENVER · PI CHUTE, CHRISTOPHER G, EICHMANN, DAVID A. · 2017 to 2021
$28.9M
The tale of two pandemics: Understanding racial and ethnic disparities from the collision of HIV and COVID-19 in the U.S.R01MH131542 · NIMH · UNIVERSITY OF WASHINGTON · PI Rena Chiman Patel · 2022 to 2026
$3.1M
NCATS NIH HHS U24 TR002306NIMH NIH HHS R01 MH131542
6 · The paper itself

Abstract

backgroundPostacute sequelae of COVID-19 (PASC), also known as long COVID, is a broad grouping of a range of long-term symptoms following acute COVID-19. These symptoms can occur across a range of biological systems, leading to challenges in determining risk factors for PASC and the causal etiology of this disorder. An understanding of characteristics that are predictive of future PASC is valuable, as this can inform the identification of high-risk individuals and future preventative efforts. However, current knowledge regarding PASC risk factors is limited.

objectiveUsing a sample of 55,257 patients (at a ratio of 1 patient with PASC to 4 matched controls) from the National COVID Cohort Collaborative, as part of the National Institutes of Health Long COVID Computational Challenge, we sought to predict individual risk of PASC diagnosis from a curated set of clinically informed covariates. The National COVID Cohort Collaborative includes electronic health records for more than 22 million patients from 84 sites across the United States.

methodsWe predicted individual PASC status, given covariate information, using Super Learner (an ensemble machine learning algorithm also known as stacking) to learn the optimal combination of gradient boosting and random forest algorithms to maximize the area under the receiver operator curve. We evaluated variable importance (Shapley values) based on 3 levels: individual features, temporal windows, and clinical domains. We externally validated these findings using a holdout set of randomly selected study sites.

resultsWe were able to predict individual PASC diagnoses accurately (area under the curve 0.874). The individual features of the length of observation period, number of health care interactions during acute COVID-19, and viral lower respiratory infection were the most predictive of subsequent PASC diagnosis. Temporally, we found that baseline characteristics were the most predictive of future PASC diagnosis, compared with characteristics immediately before, during, or after acute COVID-19. We found that the clinical domains of health care use, demographics or anthropometry, and respiratory factors were the most predictive of PASC diagnosis.

conclusionsThe methods outlined here provide an open-source, applied example of using Super Learner to predict PASC status using electronic health record data, which can be replicated across a variety of settings. Across individual predictors and clinical domains, we consistently found that factors related to health care use were the strongest predictors of PASC diagnosis. This indicates that any observational studies using PASC diagnosis as a primary outcome must rigorously account for heterogeneous health care use. Our temporal findings support the hypothesis that clinicians may be able to accurately assess the risk of PASC in patients before acute COVID-19 diagnosis, which could improve early interventions and preventive care. Our findings also highlight the importance of respiratory characteristics in PASC risk assessment. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): RR2-10.1101/2023.07.27.23293272.

Indexed as

COVID-19Post-Acute COVID-19 SyndromeAdultAgedCohort StudiesFemaleHumansMachine LearningMaleMiddle AgedRisk FactorsUnited StateschroniccovariatecovariatesCOVID-19ensembleinfectiouslong COVIDlong termmachine learningpredictpredictionpredictionspredictiverespiratoryriskrisksSARS-CoV-2sequelaestackingSuper Learner

Identifiers

PMID39146534
PMCPMC11364083

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.