Evidence map›Paper›PMID 41932876›Full record

ArticleNature communications2026

Representation learning to advance multi-institutional studies with electronic health record data from US and France.

Doudou Zhou, Han Tong, Linshanshan Wang, Suqi Liu, Xin Xiong, Ziming Gan, Griffier Romain, Boris P Hejblum, Yun-Chung Liu, Chuan Hong and 14 more

Abstract read
In one paragraph

Article in Nature communications, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

24 authors.

Doudou Zhou *Department of Statistics and Data Science, National University of Singapore, Singapore, Singapore.ORCID http://orcid.org/0000-0002-0830-2287
Han Tong *Department of Statistics, Columbia University, New York, NY, USA.ORCID http://orcid.org/0009-0002-2775-215X
Linshanshan Wang *Harvard T.H. Chan School of Public Health, Boston, MA, USA.
Suqi Liu *Harvard Medical School, Boston, MA, USA.
Xin XiongHarvard T.H. Chan School of Public Health, Boston, MA, USA.ORCID http://orcid.org/0000-0002-1162-5220
Ziming GanDepartment of Statistics, University of Chicago, Chicago, IL, USA.
Griffier RomainINSERM, Bordeaux Population Health Research Center, University Bordeaux, Bordeaux, France.
Boris P HejblumINSERM, Bordeaux Population Health Research Center, University Bordeaux, Bordeaux, France.ORCID http://orcid.org/0000-0003-0646-452X
Yun-Chung LiuDuke University, Durham, NC, USA.ORCID http://orcid.org/0000-0002-3433-5664
Chuan HongDuke University, Durham, NC, USA.
Clara-Lea BonzelHarvard T.H. Chan School of Public Health, Boston, MA, USA.
Tianrun CaiVA Boston Healthcare System, Boston, MA, USA.
Kevin PanBrown University, Providence, RI, USA.
Yuk-Lam HoVA Boston Healthcare System, Boston, MA, USA.ORCID http://orcid.org/0000-0003-3305-3830
Lauren CostaVA Boston Healthcare System, Boston, MA, USA.
Vidul A PanickanHarvard Medical School, Boston, MA, USA.ORCID http://orcid.org/0000-0003-0616-0403
J Michael GazianoHarvard Medical School, Boston, MA, USA.
Kenneth D MandlComputational Health Informatics Program, Boston Children's Hospital, Boston, MA, USA.ORCID http://orcid.org/0000-0002-9781-0477
Vianney JouhetINSERM, Bordeaux Population Health Research Center, University Bordeaux, Bordeaux, France.ORCID http://orcid.org/0000-0001-5272-2265
Rodolphe ThiebautINSERM, Bordeaux Population Health Research Center, University Bordeaux, Bordeaux, France.ORCID http://orcid.org/0000-0002-5235-3962
Zongqi XiaDepartment of Neurology, University of Pittsburgh, Pittsburgh, PA, USA.ORCID http://orcid.org/0000-0003-1500-2589
Kelly ChoHarvard Medical School, Boston, MA, USA.
Katherine LiaoHarvard Medical School, Boston, MA, USA. kliao@bwh.harvard.edu.ORCID http://orcid.org/0000-0002-4797-3200
Tianxi CaiHarvard T.H. Chan School of Public Health, Boston, MA, USA. tcai@hsph.harvard.edu.ORCID http://orcid.org/0000-0002-5379-2502

Funding

Leveraging electronic health records to optimize treatment selection and response in multiple sclerosisR01NS098023 · NINDS · UNIVERSITY OF PITTSBURGH AT PITTSBURGH · PI Zongqi Xia · 2016 to 2026
$4.6M
NINDS NIH HHS R01 NS098023
6 · The paper itself

Abstract

The widespread adoption of electronic health records has created new opportunities for translational clinical research, yet this promise remains constrained by fragmented data across privacy-siloed institutions and substantial heterogeneity in local coding practices. While privacy-preserving collaborative learning allows institutions to work together without sharing patient-level data, it does not address inconsistencies in how clinical concepts are represented across sites. We introduce a graph-based framework that addresses this gap by treating data harmonization as a scalable representation learning problem. Rather than relying on fixed standards or manual mappings, the framework integrates institution-specific summary statistics from health records, curated biomedical knowledge graphs, and semantic information derived from large language models to learn a shared semantic space. This joint learning approach aligns diverse, site-specific vocabularies while preserving patient privacy. Evaluated across seven institutions and two languages, the framework provides a robust, data-centric foundation for training and deploying clinical models across heterogeneous healthcare systems.

Indexed as

Electronic Health RecordsFranceHumansLarge Language ModelsRepresentation Machine LearningTranslational Research, BiomedicalUnited States

Identifiers

PMID41932876
PMCPMC13219506

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.