ArticleThe Lancet. Digital health2025
Re-engineering a machine learning phenotype to adapt to the changing COVID-19 landscape: a machine learning modelling study from the N3C and RECOVER consortia.
Article in The Lancet. Digital health, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 8 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
8 citing papers in PubMed.
- Characterization and validation of EHR computable phenotypes for Long COVID using patient-reported symptoms: insights from the nationwide RECOVER program.Journal of the American Medical Informatics Association : JAMIA · 2026Article
- Metformin and Severe Post-COVID-19 Outcomes Among Individuals with Diabetes Mellitus.medRxiv : the preprint server for health sciences · 2026Article
- Effect of Paxlovid treatment during acute COVID-19 on Long COVID onset: An EHR-based target trial emulation from the N3C and RECOVER consortia.PLoS medicine · 2025Article
- Long COVID Incidence Proportion in Adults and Children Between 2020 and 2024: An Electronic Health Record-Based Study From the RECOVER Initiative.Clinical infectious diseases : an official publication of the Infectious Diseases Society of America · 2025Article
- A Bayesian Survival Analysis on Long COVID and Non-Long COVID Patients: A Cohort Study Using National COVID Cohort Collaborative (N3C) Data.Bioengineering (Basel, Switzerland) · 2025Article
- Identifying commonalities and differences between EHR representations of PASC and ME/CFS in the RECOVER EHR cohort.Communications medicine · 2025Article
- EFFECT OF PAXLOVID TREATMENT DURING ACUTE COVID-19 ON LONG COVID ONSET: AN EHR-BASED TARGET TRIAL EMULATION FROM THE N3C AND RECOVER CONSORTIA.medRxiv : the preprint server for health sciences · 2025Article
- Mechanistic insights into traditional Chinese medicine for viral pneumonia treatment: signaling pathway perspectives.Frontiers in pharmacology · 2025Review
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
16 authors.
Funding
Abstract
backgroundIn 2021, we used the National COVID Cohort Collaborative (N3C) as part of the National Institutes of Health RECOVER Initiative to develop a machine learning pipeline to identify patients with a high probability of having post-acute sequelae of SARS-CoV-2 infection or long COVID. However, the increased home testing, missing documentation, and reinfections that characterise the pandemic beyond 2022 necessitated the re-engineering of our original model to account for these changes in the COVID-19 research landscape.
methodsTrained on 72 745 patient records (36 238 with long COVID and 36 507 with no evidence of long COVID), our updated XGBoost model gathered data for each patient in overlapping 100-day periods that progressed through time and issued a probability of long COVID for each 100-day period. We ran the model on patients in N3C (n=5 875 065) who met at least one of the following criteria from Jan 1, 2020, to June 22, 2023: a U07·1 (COVID-19) diagnosis code; a positive SARS-CoV-2 test; a U09·9 (post-acute sequelae of SARS-CoV-2 infection) diagnosis code; a prescription for nirmatrelvir-ritonavir or remdesivir; or an M35·81 (multisystem inflammatory syndrome in children [MIS-C]) diagnosis code. Each patient was given a model score that predicted long COVID status for each 100-day window in which they were aged ≥18 years. If a patient had known acute COVID-19 during any 100-day window (including reinfections), we censored the data from 7 days before the diagnosis or positive test date to 28 days after. We ran the model on controls selected from pre-2020 data to assess the likelihood of false positives.
findingsThe updated model had an area under the receiver operating characteristic curve of 0·90. Precision and recall could be adjusted according to a given use case, depending on whether greater sensitivity or specificity was warranted. Using our model, we estimate the overall prevalence of long COVID among the COVID-19 positive cohort within N3C repository to be 10.4%.
interpretationBy eschewing the COVID-19 index date as an anchor point for analysis, we can assess the probability of long COVID among patients who might have tested at home, or with suspected (but untested) cases of COVID-19, or multiple SARS-CoV-2 reinfections. We view this exercise as a model for maintaining and updating any machine learning pipeline used for clinical research and operations.
fundingNational Institutes of Health RECOVER Initiative.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.