Evidence map›Paper›PMID 40605780›Full record

ArticleJMIR AI2025

Harnessing Moderate-Sized Language Models for Reliable Patient Data Deidentification in Emergency Department Records: Algorithm Development, Validation, and Implementation Study.

Océane Dorémus, Dylan Russon, Benjamin Contrand, Ariel Guerra-Adames, Marta Avalos-Fernandez, Cédric Gil-Jardiné, Emmanuel Lagarde

Abstract read
In one paragraph

Article in JMIR AI, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Océane DorémusAHeaD Team, University of Bordeaux, INSERM, BPH, U1219, 146 Rue Léo Saignat, Bordeaux, F-33000, France, 33 5 57 57 15 04.ORCID http://orcid.org/0009-0000-8361-1103
Dylan RussonAHeaD Team, University of Bordeaux, INSERM, BPH, U1219, 146 Rue Léo Saignat, Bordeaux, F-33000, France, 33 5 57 57 15 04.ORCID http://orcid.org/0000-0002-1741-3560
Benjamin ContrandAHeaD Team, University of Bordeaux, INSERM, BPH, U1219, 146 Rue Léo Saignat, Bordeaux, F-33000, France, 33 5 57 57 15 04.ORCID http://orcid.org/0000-0002-2012-2676
Ariel Guerra-AdamesAHeaD Team, University of Bordeaux, INSERM, BPH, U1219, 146 Rue Léo Saignat, Bordeaux, F-33000, France, 33 5 57 57 15 04.ORCID http://orcid.org/0000-0002-7881-8246
Marta Avalos-FernandezSISTM Team, University of Bordeaux, INSERM, INRIA, BPH, U1219, Bordeaux, F-33000, France.ORCID http://orcid.org/0000-0002-5471-2615
Cédric Gil-JardinéAHeaD Team, University of Bordeaux, INSERM, BPH, U1219, 146 Rue Léo Saignat, Bordeaux, F-33000, France, 33 5 57 57 15 04.ORCID http://orcid.org/0000-0001-5329-6405
Emmanuel LagardeAHeaD Team, University of Bordeaux, INSERM, BPH, U1219, 146 Rue Léo Saignat, Bordeaux, F-33000, France, 33 5 57 57 15 04.ORCID http://orcid.org/0000-0001-8031-7400

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: The digitization of health care, facilitated by the adoption of electronic health records systems, has revolutionized data-driven medical research and patient care. While this digital transformation offers substantial benefits in health care efficiency and accessibility, it concurrently raises significant concerns over privacy and data security. Initially, the journey toward protecting patient data deidentification saw the transition from rule-based systems to more mixed approaches including machine learning for deidentifying patient data. Subsequently, the emergence of large language models has represented a further opportunity in this domain, offering unparalleled potential for enhancing the accuracy of context-sensitive deidentification. However, despite large language models offering significant potential, the deployment of the most advanced models in hospital environments is frequently hindered by data security issues and the extensive hardware resources required. Objective: The objective of our study is to design, implement, and evaluate deidentification algorithms using fine-tuned moderate-sized open-source language models, ensuring their suitability for production inference tasks on personal computers. Methods: We aimed to replace personal identifying information (PII) with generic placeholders or labeling non-PII texts as "ANONYMOUS," ensuring privacy while preserving textual integrity. Our dataset, derived from over 425,000 clinical notes from the adult emergency department of the Bordeaux University Hospital in France, underwent independent double annotation by 2 experts to create a reference for model validation with 3000 clinical notes randomly selected. Three open-source language models of manageable size were selected for their feasibility in hospital settings: Llama 2 (Meta) 7B, Mistral 7B, and Mixtral 8×7B (Mistral AI). Fine-tuning used the quantized low-rank adaptation technique. Evaluation focused on PII-level (recall, precision, and F1-score) and clinical note-level metrics (recall and BLEU [bilingual evaluation understudy] metric), assessing deidentification effectiveness and content preservation. Results: The generative model Mistral 7B performed the highest with an overall F1-score of 0.9673 (vs 0.8750 for Llama 2 and 0.8686 for Mixtral 8×7B). At the clinical notes level, the model's overall recall was 0.9326 (vs 0.6888 for Llama 2 and 0.6417 for Mixtral 8×7B). This rate increased to 0.9915 when Mistral 7B only deleted names. Four notes of 3000 failed to be fully pseudonymized for names: in 1 case, the nondeleted name belonged to a patient, while in the others, it belonged to medical staff. Beyond the fifth epoch, the BLEU score consistently exceeded 0.9864, indicating no significant text alteration. Conclusions: Our research underscores the significant capabilities of generative natural language processing models, with Mistral 7B standing out for its superior ability to deidentify clinical texts efficiently. Achieving notable performance metrics, Mistral 7B operates effectively without requiring high-end computational resources. These methods pave the way for a broader availability of pseudonymized clinical texts, enabling their use for research purposes and the optimization of the health care system.

Indexed as

clinical notesde-identificationelectronic health recordsgeneral data protection regulationlarge language modelmachine learningnatural language processingtransformers

Identifiers

PMID40605780
PMCPMC12223680

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.