ArticleJournal of biomedical informatics2025
ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis.
Article in Journal of biomedical informatics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
5 citing papers in PubMed.
- Estimate renal cell carcinoma recurrence rates using electronic health records.ESMO real world data and digital oncology · 2026Article
- Representation learning to advance multi-institutional studies with electronic health record data from US and France.Nature communications · 2026Article
- The incremental value of unstructured data via natural language processing in machine learning-based COVID-19 mortality prediction: a comparative study.BMC medical informatics and decision making · 2025Article
- Advancing the Use of Longitudinal Electronic Health Records: Tutorial for Uncovering Real-World Evidence in Chronic Disease Outcomes.Journal of medical Internet research · 2025Article
- DOME: Directional medical embedding vectors from Electronic Health Records.Journal of biomedical informatics · 2025Article
Corrections and comments
- Update of
Authors and funding
22 authors.
Funding
Abstract
objectiveElectronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. To address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features.
methodsUsing data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer's disease patients.
resultsARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.
conclusionThe proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.