Evidence map›Paper›PMID 41880603›Full record

ArticleJournal of medical Internet research2026

Improving Retrieval Augmented Generation for Health Care by Fine-Tuning Clinical Embedding Models: Development and Evaluation Study.

Kamyar Arzideh, Henning Schäfer, Ahmad Idrissi-Yaghir, Cynthia Sabrina Schmidt, Bahadir Eryilmaz, Mikel Bahn, Amin T Turki, Olivia Barbara Pollok, Eva Maria Hartmann, Philipp Winnekens and 4 more

Abstract readEvaluation Study
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

14 authors.

Kamyar ArzidehInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0009-0005-6074-804X
Henning SchäferInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0002-4123-0406
Ahmad Idrissi-YaghirInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0003-1507-9690
Cynthia Sabrina SchmidtInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0003-1994-0687
Bahadir EryilmazInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0009-0002-8743-4751
Mikel BahnInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0009-0002-0866-4023
Amin T TurkiInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0003-1347-3360
Olivia Barbara PollokInstitute of Diagnostic and Interventional Radiology and Neuroradiology, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0009-0000-0038-3986
Eva Maria HartmannInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0009-0000-2600-7217
Philipp WinnekensCentral IT Department, Data Integration Center, University Hospital Essen, Girardetstr. 2, Essen, 45131, Germany, 49 0231-77816.ORCID http://orcid.org/0009-0003-1625-3459
Katarzyna BorysInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0001-6987-6041
Johannes HauboldInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0003-4843-5911
Felix NensaInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0002-5811-7100
René HoschInstitute for Artificial Intelligence in Medicine,, University Hospital Essen, Essen, Germany.ORCID http://orcid.org/0000-0003-1760-2342

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Embedding models are critical components of Retrieval Augmented Generation (RAG) systems for retrieving and searching unstructured medical data. However, existing models are predominantly trained on publicly available English datasets, limiting their effectiveness in non-English health care settings. More importantly, these models lack training on real-world clinical documents, leading to inaccurate context retrieval when integrated into RAG systems for health care applications. This gap is particularly pronounced in specialized medical documentation containing domain-specific terminology, abbreviations, and nuanced clinical language. Objective: This retrospective study aimed to develop and validate embedding models specifically trained on real-world clinical documents from multiple medical specialties to improve medical information retrieval (IR) and RAG system performance in both German and English language contexts. Methods: We fine-tuned embedding models, so-called sentence transformers, using the multilingual-e5-large architecture as a foundation. Training data consisted of approximately 11 million question-answer pairs synthetically generated from 400,000 diverse clinical documents from a large German tertiary hospital, spanning 163,840 patients and 282,728 clinical cases between 2018 and 2023. The SauerkrautLM-SOLAR-Instruct large language model generated medically relevant questions and corresponding answers for each document. The dataset was additionally pseudonymized and translated into English to aim for broader applicability. Models were evaluated in 2 distinct scenarios: IR using questions with multiple relevant passages, and RAG system performance in both cross-patient and patient-centered contexts. Results: In the IR evaluation, the fine-tuned miracle model achieved a mAP@100 of 0.27, outperforming the multilingual-e5-large baseline (0.14) and state-of-the-art models such as bge-m3 (0.11). In the RAG evaluation, the model demonstrated robust performance comparable with the baseline in the constrained patient-centered scenario (BERTScore F1 0.781 vs 0.778) and showed moderate improvements in the unconstrained cross-patient setting (BLEURT 0.56 vs 0.53). Notably, the model trained on pseudonymized data achieved comparable retrieval performance (mAP@100 0.25) and the highest scores for patient-centered contextual precision (0.93). Performance gains were robust in the German dataset, while the translated English model demonstrated promising results as a proof of concept for cross-lingual transfer. Conclusions: By leveraging a comprehensive real-world dataset spanning multiple medical specialties and using large language models for synthetic question generation, we successfully created and validated domain-specific embedding models. These models can improve medical IR in large-scale search spaces and perform competitively in constrained RAG applications. By publishing the models trained on pseudonymized data, other health care institutions can integrate or adapt these embedding models to their needs. This work establishes a reproducible framework for developing domain-specific clinical embedding models, with the potential to improve data retrieval in medical settings.

Indexed as

Delivery of Health CareElectronic Health RecordsInformation Storage and RetrievalHumansLarge Language ModelsNatural Language ProcessingRetrospective Studiesinformation retrievallarge language modelsLLMnatural language processingNLPRAGRetrieval Augmented Generation

Identifiers

PMID41880603
PMCPMC13016438

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.