ArticleNPJ digital medicine2025
Iterative refinement and goal articulation to optimize large language models for clinical information extraction.
Article in NPJ digital medicine, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 16 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
16 citing papers in PubMed.
- Benchmarking commercial large language models for gene-disease-phenotype extraction from full-text human genetics literature.Quantitative biology (Beijing, China) · 2026Article
- Automated Identification of Complex Percutaneous Coronary Intervention from Cardiac Catheterization Reports using Large Language Models.medRxiv : the preprint server for health sciences · 2026Article
- Platform-Mediated Visibility of Exploitation Indicators: A Cross-Platform Analysis of AdultWork (UK) and CallEscort (US).Behavioral sciences (Basel, Switzerland) · 2026Article
- Hierarchical multi-label structuring of Japanese SOAP clinical notes with large language models.Scientific reports · 2026Article
- OntoCodex: a multi-agent biomedical ontology enrichment framework.npj health systems · 2026Article
- RelAgent: a multi-agent solution for molecular relationship grounding.Bioinformatics (Oxford, England) · 2026Article
- Scalable extraction of social determinants of health from clinical notes in a sepsis cohort using instruction-tuned language models.JAMIA open · 2026Article
- clickBrick prompt engineering: optimizing large language model performance in clinical psychiatry.Npj mental health research · 2026Article
- Clinical large language model centered on electronic medical records.NPJ digital medicine · 2026Article
- Performance of large language models for extracting clinical data from breast cancer pathology reports: a systematic review.NPJ digital medicine · 2026Article
- Artificial Intelligence Applications for Automated Data Extraction and Secondary Use of Clinical Information in Uro-oncology: A Systematic Review.European urology open science · 2026Review
- Hybrid rule-based and on-premises LLM pipeline for extracting CMR and CPET metrics from free-text reports in repaired tetralogy of Fallot.medRxiv : the preprint server for health sciences · 2026Article
- Precision oncology: from large language models to multi-agent systems.Frontiers in oncology · 2026Review
- Radiologic, Pathologic, and Deep Learning Predictors of Response to Immune Checkpoint Blockade in Renal Cell Carcinoma Patients Undergoing Post-Treatment Nephrectomy.medRxiv : the preprint server for health sciences · 2025Article
- Large Language Models in Population Oncology: A Contemporary Review on the Use of Large Language Models to Support Data Collection, Aggregation, and Analysis in Cancer Care and Research.JCO clinical cancer informatics · 2025Review
- From Mutation to Prognosis: AI-HOPE-PI3K Enables Artificial Intelligence Agent-Driven Integration of PI3K Pathway Data in Colorectal Cancer Precision Medicine.International journal of molecular sciences · 2025Article
Corrections and comments
- Update of
Authors and funding
13 authors.
Funding
Abstract
Extracting structured data from free-text medical records at scale is laborious, and traditional approaches struggle in complex clinical domains. We present a novel, end-to-end pipeline leveraging large language models (LLMs) for highly accurate information extraction and normalization from unstructured pathology reports, focusing initially on kidney tumors. Our innovation combines flexible prompt templates, the direct production of analysis-ready tabular data, and a rigorous, human-in-the-loop iterative refinement process guided by a comprehensive error ontology. Applying the finalized pipeline to 2297 kidney tumor reports with pre-existing templated data available for validation yielded a macro-averaged F1 of 0.99 for six kidney tumor subtypes and 0.97 for detecting kidney metastasis. We further demonstrate flexibility with multiple LLM backbones and adaptability to new domains, utilizing publicly available breast and prostate cancer reports. Beyond performance metrics or pipeline specifics, we emphasize the critical importance of task definition, interdisciplinary collaboration, and complexity management in LLM-based clinical workflows.
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.