Evidence map›Paper›PMID 42770666›Full record

ArticleJournal of medical Internet research2026

Improving Reliability and Explainability of Medical Question Answering Through Atomic Fact-Checking in Retrieval-Augmented Large Language Models: Creation and Validation Study.

Juraj Vladika, Annika Domres, Mai Nguyen, Rebecca Moser, Jana Nano, Felix Busch, Lisa Adams, Keno K Bressem, Denise Bernhardt, Stephanie E Combs and 3 more

Abstract readValidation Study
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. A Scoping Review of Large Language Models in Personal Sleep Wellness.Mayo Clinic proceedings. Digital health · 2025
    Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

13 authors.

Juraj Vladika *Department of Computer Science, TUM School of Computation Information and Technology, Technical University of Munich, Garching, Bavaria, Germany.
Annika Domres *Department of Radiation Oncology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Ismaninger Str. 22, Munich, Bavaria, Germany, 49 089 4140-4501.ORCID http://orcid.org/0009-0009-8822-2620
Mai NguyenDepartment of Radiation Oncology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Ismaninger Str. 22, Munich, Bavaria, Germany, 49 089 4140-4501.
Rebecca MoserDepartment of Radiation Oncology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Ismaninger Str. 22, Munich, Bavaria, Germany, 49 089 4140-4501.
Jana NanoDepartment of Radiation Oncology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Ismaninger Str. 22, Munich, Bavaria, Germany, 49 089 4140-4501.
Felix BuschDepartment of Diagnostic and Interventional Radiology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Munich, Germany.ORCID http://orcid.org/0000-0001-9770-8555
Lisa AdamsDepartment of Diagnostic and Interventional Radiology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Munich, Germany.ORCID http://orcid.org/0000-0001-5836-4542
Keno K BressemDepartment of Diagnostic and Interventional Radiology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Munich, Germany.ORCID http://orcid.org/0000-0001-9249-8624
Denise BernhardtDepartment of Radiation Oncology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Ismaninger Str. 22, Munich, Bavaria, Germany, 49 089 4140-4501.
Stephanie E CombsDepartment of Radiation Oncology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Ismaninger Str. 22, Munich, Bavaria, Germany, 49 089 4140-4501.
Kai BormDepartment of Radiation Oncology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Ismaninger Str. 22, Munich, Bavaria, Germany, 49 089 4140-4501.
Florian MatthesDepartment of Computer Science, TUM School of Computation Information and Technology, Technical University of Munich, Garching, Bavaria, Germany.ORCID http://orcid.org/0000-0002-6667-5452
Jan C PeekenDepartment of Radiation Oncology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Ismaninger Str. 22, Munich, Bavaria, Germany, 49 089 4140-4501.ORCID http://orcid.org/0000-0003-2679-9853

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Large language models (LLMs) exhibit extensive medical knowledge but are prone to hallucinations and show low fact-level explainability, limiting clinical adoption and regulatory compliance. Existing approaches, such as retrieval-augmented generation, partially address these issues by grounding answers in source documents; however, the aforementioned problems persist. Objective: We propose the application of an atomic fact-checking framework designed to enhance the reliability and explainability of LLMs in medical long-form question answering. By decomposing generated answers into discrete atomic facts and verifying each against an authoritative knowledge base of medical guidelines, this approach enables precise identification and correction of incorrect statements, alongside explicit linkage to supporting literature. Methods: The fact-checking algorithm operates within a retrieval-augmented generation framework: LLM-generated answers are decomposed into atomic facts (smallest and self-contained information units), each of which is assessed and corrected if FALSE. To determine an optimal strategy, the validation-question and answer (Q&A) set on prostate cancer treatment was tested under varying instructions. An extensive evaluation, including multireader assessments by human medical experts and the automated open Q&A benchmark AMEGA (Autonomous Medical Evaluation for Guideline Adherence), was conducted for the final pipeline. In addition to another radiooncologic test-Q&A set, anonymized real-world tumor board cases and an independent, established neurology-Q&A set were used. Given their transparency and accessibility advantages, we compared various open-source models in pairs of generalist models and their medical fine-tuned counterparts, with regard to performance and improvements by fact-checking. Results: The framework significantly reduced hallucinations and inaccuracies. Medical expert assessment and automated benchmarks demonstrated significant improvements in factual accuracy, achieving up to a 50% overall answer improvement and an 80% hallucination detection rate. Notably, the observed gain was strongest in real tumor-board questions-the most challenging dataset. Additionally, the framework achieved high explainability by tracing each atomic fact back to the most relevant chunks from the database, providing a granular, transparent explanation of the generated responses. Conclusions: To conclude, we present the application of an atomic fact-checking algorithm to medical Q&A. It identifies factual inaccuracies and hallucinations in LLM-generated answers, achieving the greatest gains on clinically realistic, complex questions. Correction via fact-checking improves the overall answer quality while achieving fact-wise explainability, paving the way for more credible clinical use of LLMs.

Indexed as

Information Storage and RetrievalLarge Language ModelsAlgorithmsHumansMaleProstatic NeoplasmsReproducibility of Resultsatomic factatomic fact-checkingautoevaluationbacktracingfact-checkinghallucinationlarge language modelLLMmedical fine-tunedmedical Q&Aopen-sourceprompt engineeringquestion and answerradiation oncologyRAGretrieval-augmented generationrubrics

Identifiers

PMID42770666
PMCPMC13595421

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.