Evidence map›Paper›PMID 42686197›Full record

ArticleJournal of medical Internet research2026

A Secure, Scalable Large Language Model-Based System (CIDER) for High-Throughput Clinical Data Extraction From Medical Reports: Retrospective Validation Study.

Máté Posta, Aida Figler, Zsófia Dobolyi, Balázs Győrffy

Abstract readValidation Study
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Máté PostaInstitute of Molecular Life Sciences, HUN-REN Research Centre for Natural Sciences, Budapest, Hungary.ORCID https://orcid.org/0009-0002-4912-3848
Aida FiglerDepartment of Bioinformatics, Semmelweis University, Budapest, Hungary.ORCID https://orcid.org/0009-0004-1279-6226
Zsófia DobolyiDepartment of Bioinformatics, Semmelweis University, Budapest, Hungary.ORCID https://orcid.org/0009-0003-1563-1550
Balázs GyőrffyInstitute of Molecular Life Sciences, HUN-REN Research Centre for Natural Sciences, Budapest, Hungary.ORCID https://orcid.org/0000-0002-5772-3766

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundA substantial proportion of clinically relevant information remains locked in unstructured narrative documents, creating a bottleneck for clinical research, biobank annotation, registry development, and real-world evidence generation. While large language models (LLMs) enable advanced clinical text mining, adoption is constrained by concerns regarding data security, multilingual performance, and reproducibility. Manual data abstraction remains predominant for registry curation and retrospective research, despite being labor intensive, costly, and prone to variability.

objectiveWe developed and validated CIDER (Clinical Data Extractor), a secure, institutionally deployable, LLM-based pipeline for automated structured data extraction from routine clinical reports. We assessed the potential utility of the system for improving the completeness of clinical research datasets.

methodsCIDER uses an asynchronous FastAPI-based architecture with a locally deployed vLLM inference engine running Qwen3-VL-32B-Instruct-FP8 model in an institution-controlled environment. The system was validated on 2073 real-world Hungarian-language histopathology reports (a challenging non-English setting), using a manually curated structured database as the reference standard. Seven variables were evaluated (sex, surgery year, T stage, N stage, organ, histology, and size). Extraction performance was assessed using exact-match accuracy, weighted F

resultsThe validation dataset comprised stand-alone native-text PDF pathology reports originating from multiple Hungarian oncology centers. Input document length showed a median of 3926 (mean 4258, SD 1057, IQR 3490-4688) tokens, while generated outputs contained a median of 63 (mean 62.5, SD 8.2, IQR 56-66) tokens. At a temperature of 0.1, CIDER achieved near-human agreement with expert-curated reference database, with exact-match accuracies of 99.5% for sex, 98.1% for surgery year, 95.8% for organ, 95.6% for T stage, 92.4% for N stage, 87.5% for histology, and 78.1% for tumor size. Weighted F

conclusionsCIDER demonstrates that locally deployed open-weight LLMs can reliably extract structured clinical data from complex pathology reports while preserving institutional control over sensitive data. These findings support the feasibility of secure, institutionally deployable, LLM-based extraction systems for generating research-ready datasets, facilitating clinical registry development, improving dataset completeness, and enabling scalable reuse of unstructured clinical documentation.

Indexed as

Data MiningInformation Storage and RetrievalHumansLarge Language ModelsReproducibility of ResultsRetrospective Studiesbiobankingfree-text documentshealth care recordshistopathologylarge language modelLLMnatural language processingNLPoncologypatient documentation

Identifiers

PMID42686197
PMCPMC13583534

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.