Evidence map›Paper›PMID 42783862›Full record

ArticleJournal of imaging2026

A Multimodal AI Framework for Medical Education: Integrating Adaptive Image Retrieval, Fast Synthesis, and LLM-Based Clinical Auditing.

Miguel Díaz-Benito, Cecilia Diana-Albelda, Álvaro García-Martín, Mario Rubén Paz Campos, Jesus Bescos

Abstract read
In one paragraph

Article in Journal of imaging, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Miguel Díaz-BenitoDepartment of Signal Theory and Communications, Universidad Carlos III de Madrid, 28911 Leganés, Spain.ORCID 0009-0002-2110-9972
Cecilia Diana-AlbeldaVideo Processing and Understanding Lab, Escuela Politécnica Superior, Universidad Autónoma de Madrid, 28049 Madrid, Spain.ORCID 0009-0009-9210-0853
Álvaro García-MartínVideo Processing and Understanding Lab, Escuela Politécnica Superior, Universidad Autónoma de Madrid, 28049 Madrid, Spain.ORCID 0000-0002-1705-3972
Mario Rubén Paz CamposHospital Universitario de La Princesa, 28006 Madrid, Spain.ORCID 0009-0007-4964-5626
Jesus BescosVideo Processing and Understanding Lab, Escuela Politécnica Superior, Universidad Autónoma de Madrid, 28049 Madrid, Spain.ORCID 0000-0001-6238-6859

Funding

Ministerio de Ciencia e Innovación of the Spanish Government PID2021-125051OB-I00Regional Government of Madrid of Spain TEC 2024/COM-322
6 · The paper itself

Abstract

Access to reliable medical images is essential for clinical training. To address this need, this paper presents an extended version of MIRAGE, a multimodal retrieval and generation system that utilizes a shared latent space to process medical queries by retrieving real images from the ROCO dataset, generating synthetic scans, and providing LLM-based clinical descriptions alongside dual-concept visual comparisons. To overcome previous computational limits and the lack of clinical validation, we introduce three core enhancements: first, an Auto-α module to dynamically weight visual and textual similarities; second, the integration of LCM-LoRA to accelerate synthetic image generation; and third, an automated clinical auditor based on Gemini 2.5 Flash. Experimental results demonstrate that Auto-α improves retrieval accuracy for heterogeneous queries, reaching 38.83% Top-1 Recall over a 65,419-image gallery and outperforming nine fusion baselines evaluated under a unified configuration, with a controlled ablation attributing most of this gain to learning the weight rather than merely making it query-adaptive, while the LCM-LoRA module reduces computational costs by a factor of 12.5× in CPU environments, with a blinded radiologist evaluation confirming only a small drop in clinical quality. Furthermore, the clinical auditor achieves a 0.805 Pearson correlation against an expert radiologist, effectively correcting the systematic overestimation of traditional CLIP scores. Finally, the optimized platform is publicly deployed on Hugging Face.

Indexed as

Auto-αclinical auditdiffusion modelsimage retrievalLCM-LoRAmedical educationmultimodal language models

Identifiers

PMID42783862
PMCPMC13608127

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.