Evidence map›Paper›PMID 42368617›Full record

ArticleEULAR rheumatology open2025

Optimising the clinical application of rheumatology guidelines using large language models: a retrieval-augmented generation framework integrating EULAR and ACR recommendations.

Alfredo Madrid-García, Diego Benavent, Chamaida Plasencia-Rodríguez, Zulema Rosales-Rosado, Beatriz Merino-Barbancho, Dalifer Freites-Núñez

Abstract read
In one paragraph

Article in EULAR rheumatology open, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers.

0numbers the graph read from it
0cells of the map it votes in
6citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

6 citing papers in PubMed.

  1. Review
  2. Article
  3. Article
  4. Article
  5. Article
  6. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Alfredo Madrid-GarcíaIndependent Researcher.
Diego BenaventRheumatology Department, Hospital Universitari de Bellvitge, Barcelona, Spain.
Chamaida Plasencia-RodríguezRheumatology Department, Hospital Universitario La Paz-IdiPaz, Madrid, Spain.
Zulema Rosales-RosadoGrupo de Patología Musculoesquelética. Hospital Clínico San Carlos. Instituto de Investigación Sanitaria San Carlos (IdISSC), Madrid, Spain.
Beatriz Merino-BarbanchoEscuela Técnica Superior de Ingenieros de Telecomunicación. Universidad Politécnica de Madrid. Avenida Complutense, Madrid, Spain.
Dalifer Freites-NúñezGrupo de Patología Musculoesquelética. Hospital Clínico San Carlos. Instituto de Investigación Sanitaria San Carlos (IdISSC), Madrid, Spain.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objectives: Timely access to current rheumatology guidelines at the point of care is challenging. We aimed to develop and evaluate the first retrieval-augmented generation (RAG) system designed for adult rheumatology, integrating European Alliance of Associations for Rheumatology (EULAR) and American College of Rheumatology (ACR) guidelines to provide rheumatologists with timely evidence-based recommendations. Methods: EULAR and ACR management guidelines were selected by rheumatologists based on their clinical relevance for decision-making and processed. A RAG system was implemented. To evaluate it, 10 questions per guideline were generated using ChatGPT 4.5. Answers to these were produced by ChatGPT-o3-mini with context retrieval (RAG) and without (baseline). Performance was assessed by an Large language model (LLM)-as-a-judge (Gemini 2.0 Flash) using a 5-point Likert scale across 5 dimensions: relevance, factual accuracy, safety, completeness, and conciseness; it also determined preference between the RAG and baseline responses. For validation, 2 rheumatologists independently evaluated a random sample of questions (15%) on the same domains. Statistical significance was established using the Wilcoxon signed-rank and binomial tests. Results: Seventy-four guidelines were included, yielding 740 evaluation questions. The LLM-as-a-judge evaluation showed the RAG system significantly outperformed the baseline across all criteria ( Conclusions: Developing a RAG system integrating extensive EULAR/ACR rheumatology guidelines improves answer quality compared to a baseline LLM. This evaluation provides a robust foundation for reliable, artificial intelligence-driven clinical decision support tools designed to enhance evidence-based practice by providing clinicians with rapid, context-aware access to recommendations.

Identifiers

PMID42368617
PMCPMC13292420

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.