ArticleESMO real world data and digital oncology2026
Investigating fine-tuning versus zero-shot learning for general large language models when predicting cancer survival from initial oncology consultation documents.
Article in ESMO real world data and digital oncology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
8 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Unstructured oncology consultation notes contain rich clinical information that may support survival prediction. Open-weight large language models (LLMs) can utilize these notes with zero-shot inference or fine-tuning, but their relative value for this setting remains unclear. The objective of this study is to evaluate open-weight LLMs for predicting 60-month survival from initial oncology consultation notes, comparing (i) zero-shot performance, (ii) performance after fine-tuning, and (iii) smaller natural language processing models trained on the same dataset in prior work. Materials and methods: We used Meta's Llama models to predict patients' 60-month survival using oncology consultation notes from a dataset of 59 800 patients. We tested both zero-shot and fine-tuning approaches. Metrics included balanced accuracy (BA) and weighted F1. Results: Zero-shot performance was limited. Llama-2-13B performed best among the zero-shot configurations (average performance across prompts: BA 0.596, weighted F1 0.644; performance on Prompt 4: BA 0.766, weighted F1 0.802). Fine-tuning improved performance across models: Llama-2-13B achieved BA 0.842, weighted F1 0.846, area under the receiver operating characteristic curve (AUC) 0.905; Llama-2-7B achieved BA 0.840, weighted F1 0.843, AUC 0.911; Llama-3.1-8B achieved BA 0.829, weighted F1 0.829, AUC 0.881. Performance was numerically similar to smaller models trained on the same task and data. Conclusions: For predicting 60-month survival from initial oncology consultation documents, fine-tuning open-weight LLMs meaningfully improves performance compared with zero-shot use, but does not consistently outperform smaller language models. This may suggest that both fine-tuned LLMs and smaller models merit continued investigation, with the most appropriate approach likely to depend on the outcome of interest, clinical context, and practical considerations such as hardware, privacy, and deployment feasibility.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.