Evidence map›Paper›PMID 42825298›Full record

ArticleJAMIA open2026

Harmonizing UK primary care prescription records for research: a case study in the UK biobank.

Cai R Ytsma, Ana Torralbo, Natalie K Fitzpatrick, Maik Pietzner, Daniela Nguyen, Ioannis Louloudis, Sabrina Ansarey, Spiros Denaxas

Abstract read
In one paragraph

Article in JAMIA open, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Cai R YtsmaInstitute of Health Informatics, University College London, London, NW1 2DA, United Kingdom.ORCID https://orcid.org/0000-0002-8333-6428
Ana TorralboInstitute of Health Informatics, University College London, London, NW1 2DA, United Kingdom.
Natalie K FitzpatrickInstitute of Health Informatics, University College London, London, NW1 2DA, United Kingdom.
Maik PietznerHealth Data Modelling, Berlin Institute of Health, Berlin, 10117, Germany.
Daniela NguyenPrecision Healthcare University Research Institute, Queen Mary University of London, London, E1 1HH, United Kingdom.
Ioannis LouloudisNovo Nordisk Foundation Center for Protein Research, University of Copenhagen, Copenhagen, 2200, Denmark.
Sabrina AnsareyInstitute of Health Informatics, University College London, London, NW1 2DA, United Kingdom.
Spiros DenaxasInstitute of Health Informatics, University College London, London, NW1 2DA, United Kingdom.ORCID https://orcid.org/0000-0001-9612-7791

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objective: We aimed to develop and validate an automated, scalable framework to harmonize fragmented UK primary care prescription records into a research-ready dataset by mapping diverse medical ontologies to a unified reference standard from the UK Biobank. Materials and methods: Preprocessing involved selecting a single drug code where multiple were recorded, cleaning codes to match reference presentations, expanding code granularity via drug descriptions, and updating outdated codes. Harmonization mapped British National Formulary (BNF) and Read2 codes to dm+d, the National Health Service standard vocabulary, with records then homogenized to the Virtual Medicinal Product (VMP) granularity. We validated our approach by profiling prescribing patterns across 312 diseases. Results: We preprocessed 57 659 844 records from 221 868 participants; we dropped 48 950 records for missing drug codes and 13% used multiple ontologies. Most records were BNF-encoded (76%), and 49% required granularity expansion via drug description. In total, 72% of records were harmonized to dm+d, of which 99.98% were converted to VMP. Across 312 diseases, we identified 23 352 disease-drug associations involving 237 medications (BNF subparagraphs) that survived statistical correction, most resembling drug-indication pairs. Discussion: Our framework addresses a longstanding barrier to large-scale pharmacoepidemiological research by reconciling heterogeneous coding systems into a single, consistent standard. The high conversion rate to VMP demonstrates near-complete coverage, while the recovery of plausible drug-indication associations supports validity. Limitations include residual uncodable records and dependence on dm+d completeness. Conclusion: Our approach transforms fragmented prescription records into a streamlined, enriched dataset suitable for large-scale discovery research.

Indexed as

data science (D000077488)electronic health records (D057286)epidemiologic studies (D016021)medical coding (D059019)prescription drugs (D055553)

Identifiers

PMID42825298
PMCPMC13630483

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.