ArticleNPJ digital medicine2025
Leveraging long context in retrieval augmented language models for medical question answering.
Article in NPJ digital medicine, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 31 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
31 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Clinical applications of large language models in knee osteoarthritis: a systematic review.Frontiers in medicine · 2025Pooled it
- Generative large language models in medicine: a scoping review of recent methodological advances.npj health systems · 2026Review
- Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Health Care: Scoping Review.Journal of medical Internet research · 2026Article
- Evaluating the First CE-Marked LLM-Based Chatbot for Evidence-Based Neurology.European journal of neurology · 2026Article
- A Multiagent Large Language Model Framework for Emergency Treatment Recommendation in Acute Ischemic Stroke: Development and Validation Study.Journal of medical Internet research · 2026Article
- Initial-Visit Specialty Triage in Rare Diseases Using Large Language Models: Retrospective Benchmarking Study.Journal of medical Internet research · 2026Article
- Improving Retrieval-Augmented Generation without Taxonomy-based Error Categorization.Proceedings of the conference. Association for Computational Linguistics. Meeting · 2026Article
- Potential and limitations of a large language model in the statistical review of comparative categorical data: an exploratory study of structured prompt-guided approach.Research integrity and peer review · 2026Article
- FaithfulnessFindings of ACL. ACL · 2026Article
- Precision Grounding: augmenting large language models with evidence-based databases for trustworthy genetic variant summarization.International journal of medical informatics · 2026Article
- Clinical large language model centered on electronic medical records.NPJ digital medicine · 2026Article
- Optimizing GPT-5 for Operation-Procedure-Code-Extraction from Operative Reports in Meningioma Surgery: Feasibility and Comparison of Context-Enhancements.Applied clinical informatics · 2026Article
- Exploring Nurses' Perspectives on the Use of Artificial Intelligence Chatbots for Mental Health Support: A Cross-Sectional Study in Greece.Nursing reports (Pavia, Italy) · 2026Article
- A neural-symbolic AI agent system for biomedical concept mapping.NPJ digital medicine · 2026Article
- Improving support and self-management of ophthalmic patients using an artificial intelligence health coach.Indian journal of ophthalmology · 2026Article
- Retrieval‑augmented large language models for depression screening and suicide risk stratification.BMC psychiatry · 2026Article
- Modeling discourse structure with 2D similarity-based random walks for improved understanding of online conversations.Scientific reports · 2026Article
- Benchmarking Large Language Models for Drug Combination Alerts: Achieving Expert-Level Reliability via Knowledge Grounding and Contextual Reasoning.Journal of medicinal chemistry · 2026Article
- Multi-Evidence Clinical Reasoning With Retrieval-Augmented Generation for Emergency Triage: Retrospective Evaluation Study.JMIR medical informatics · 2026Article
- Federated knowledge retrieval elevates large language model performance on biomedical benchmarks.GigaScience · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
11 authors.
Funding
Abstract
While holding great promise for improving and facilitating healthcare through applications of medical literature summarization, large language models (LLMs) struggle to produce up-to-date responses on evolving topics due to outdated knowledge or hallucination. Retrieval-augmented generation (RAG) is a pivotal innovation that improves the accuracy and relevance of LLM responses by integrating LLMs with a search engine and external sources of knowledge. However, the quality of RAG responses can be largely impacted by the rank and density of key information in the retrieval results, such as the "lost-in-the-middle" problem. In this work, we aim to improve the robustness and reliability of the RAG workflow in the medical domain. Specifically, we propose a map-reduce strategy, BriefContext, to combat the "lost-in-the-middle" issue without modifying the model weights. We demonstrated the advantage of the workflow with various LLM backbones and on multiple QA datasets. This method promises to improve the safety and reliability of LLMs deployed in healthcare domains by reducing the risk of misinformation, ensuring critical clinical content is retained in generated responses, and enabling more trustworthy use of LLMs in critical tasks such as medical question answering, clinical decision support, and patient-facing applications.
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.