ArticleJournal of medical Internet research2026
A Secure, Scalable Large Language Model-Based System (CIDER) for High-Throughput Clinical Data Extraction From Medical Reports: Retrospective Validation Study.
Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundA substantial proportion of clinically relevant information remains locked in unstructured narrative documents, creating a bottleneck for clinical research, biobank annotation, registry development, and real-world evidence generation. While large language models (LLMs) enable advanced clinical text mining, adoption is constrained by concerns regarding data security, multilingual performance, and reproducibility. Manual data abstraction remains predominant for registry curation and retrospective research, despite being labor intensive, costly, and prone to variability.
objectiveWe developed and validated CIDER (Clinical Data Extractor), a secure, institutionally deployable, LLM-based pipeline for automated structured data extraction from routine clinical reports. We assessed the potential utility of the system for improving the completeness of clinical research datasets.
methodsCIDER uses an asynchronous FastAPI-based architecture with a locally deployed vLLM inference engine running Qwen3-VL-32B-Instruct-FP8 model in an institution-controlled environment. The system was validated on 2073 real-world Hungarian-language histopathology reports (a challenging non-English setting), using a manually curated structured database as the reference standard. Seven variables were evaluated (sex, surgery year, T stage, N stage, organ, histology, and size). Extraction performance was assessed using exact-match accuracy, weighted F
resultsThe validation dataset comprised stand-alone native-text PDF pathology reports originating from multiple Hungarian oncology centers. Input document length showed a median of 3926 (mean 4258, SD 1057, IQR 3490-4688) tokens, while generated outputs contained a median of 63 (mean 62.5, SD 8.2, IQR 56-66) tokens. At a temperature of 0.1, CIDER achieved near-human agreement with expert-curated reference database, with exact-match accuracies of 99.5% for sex, 98.1% for surgery year, 95.8% for organ, 95.6% for T stage, 92.4% for N stage, 87.5% for histology, and 78.1% for tumor size. Weighted F
conclusionsCIDER demonstrates that locally deployed open-weight LLMs can reliably extract structured clinical data from complex pathology reports while preserving institutional control over sensitive data. These findings support the feasibility of secure, institutionally deployable, LLM-based extraction systems for generating research-ready datasets, facilitating clinical registry development, improving dataset completeness, and enabling scalable reuse of unstructured clinical documentation.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.