Evidence map›Paper›PMID 41982780›Full record

ArticleJournal of the Royal Statistical Society. Series B, Statistical methodology2026

Inference of dependency knowledge graph for Electronic Health Records.

Zhiwei Xu, Ziming Gan, Doudou Zhou, Shuting Shen, Junwei Lu, Tianxi Cai

Erratum issuedAbstract read
In one paragraph

Article in Journal of the Royal Statistical Society. Series B, Statistical methodology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. An erratum has been issued. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

6 authors.

Zhiwei XuDepartment of Statistics, University of Michigan, Ann Arbor, MI 48109, USA.
Ziming GanDepartment of Statistics, University of Chicago, Chicago, Il 60637, USA.
Doudou ZhouDepartment of Statistics and Data Science, National University of Singapore, Singapore 119077, Singapore.ORCID https://orcid.org/0000-0002-0830-2287
Shuting ShenDepartment of Statistics and Data Science, National University of Singapore, Singapore 119077, Singapore.
Junwei LuDepartment of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA 02115, USA.
Tianxi CaiDepartment of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA 02115, USA.ORCID https://orcid.org/0000-0002-5379-2502

Funding

Leveraging electronic health records to optimize treatment selection and response in multiple sclerosisR01NS098023 · NINDS · UNIVERSITY OF PITTSBURGH AT PITTSBURGH · PI Zongqi Xia · 2016 to 2026
$4.6M
Semi-supervised Approaches to Denoising Electronic Health Records Data for Risk PredictionR01LM013614 · NLM · HARVARD UNIVERSITY D/B/A HARVARD SCHOOL OF PUBLIC HEALTH · PI CAI, TIANXI, GUO, ZIJIAN · 2021 to 2024
$1.4M
NINDS NIH HHS R01 NS098023NLM NIH HHS R01 LM013614
6 · The paper itself

Abstract

The effective analysis of high-dimensional Electronic Health Record (EHR) data, with substantial potential for healthcare research, presents notable methodological challenges. Employing predictive modeling guided by a knowledge graph (KG), which enables efficient feature selection, can enhance both statistical efficiency and interpretability. While various methods have emerged for constructing KGs, existing techniques often lack statistical certainty concerning the presence of links between entities, especially in scenarios where the utilization of patient-level EHR data is limited due to privacy concerns. In this paper, we propose the first inferential framework for deriving a sparse KG with statistical guarantee based on a dynamic log-linear topic model. Within this model, the KG embeddings are estimated by performing singular value decomposition on the empirical pointwise mutual information matrix, offering a scalable solution. We then establish entrywise asymptotic normality for the KG low-rank estimator, enabling the recovery of sparse graph edges with controlled type I error. Our work uniquely addresses the under-explored domain of statistical inference about non-linear statistics under the low-rank temporal dependent models, a critical gap in existing research. We validate our approach through extensive simulation studies and then apply the method to real-world EHR data in constructing clinical KGs and generating clinical feature embeddings.

Indexed as

hypothesis testingknowledge graph embeddinglow-rank modelsnon-linear structure

Identifiers

PMID41982780
PMCPMC13070795

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.