Evidence map›Paper›PMID 42057294›Full record

ArticleBioinformatics (Oxford, England)2026

Metappuccino: large language model-driven reconstruction of sequence read archive metadata for cancer research.

Fiona Hak, Camille Marchet, Daniel Gautheret, Mélina Gallopin

Abstract read
In one paragraph

Article in Bioinformatics (Oxford, England), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Fiona HakUniversité Paris-Saclay, CEA, CNRS, Institute for Integrative Biology of the Cell (I2BC), Gif-sur-Yvette 91190, France.ORCID 0009-0002-8296-1747
Camille MarchetUniversité de Lille, CNRS, CRIStAL UMR 9189, Lille F-59000, France.ORCID 0000-0002-7235-7346
Daniel GautheretUniversité Paris-Saclay, CEA, CNRS, Institute for Integrative Biology of the Cell (I2BC), Gif-sur-Yvette 91190, France.ORCID 0000-0002-1508-8469
Mélina GallopinUniversité Paris-Saclay, CEA, CNRS, Institute for Integrative Biology of the Cell (I2BC), Gif-sur-Yvette 91190, France.ORCID 0000-0002-2431-7825

Funding

2021-2030 Cancer Control StrategyAgence Nationale de la Recherche ANR-22-CE45-0007
6 · The paper itself

Abstract

motivationHigh-throughput RNA sequencing has significantly advanced transcriptomic profiling in oncology. Millions of RNA-seq datasets have accumulated in public databases such as the Sequence Read Archive (SRA). However, fragmented, ambiguous, or missing metadata can severely limit accurate cohort selection, introduce bias, and delay discoveries.

resultsTo address these issues, we introduce 'Metappuccino', a hybrid metadata enrichment tool built on Mistral-7B-Instruct and specialized via low-rank adaptation (LoRA). Metappuccino reconstructs 19 metadata classes (e.g. organ, disease, cell type) by combining deterministic extraction/normalization with model-based completion: 4 submission-mandatory fields are read directly from SRA/API records, while the remaining 15 classes are obtained through validated rule-based extraction when explicitly supported by the context and otherwise predicted by the LoRA-specialized model when information is missing or ambiguous. To promote robust, context-aware inference rather than memorization, we designed training and data partitioning to minimize leakage and preserve generalization. When applicable, predicted values are mapped to standardized ontologies to ensure consistent, interoperable annotations. Across our benchmarks, Metappuccino substantially improves accuracy over the base model, matches or exceeds recent larger open-source LLMs, and reduces inference time by up to two-fold relative to these baselines. By enriching under-annotated public RNA-seq records, Metappuccino increases the usability of SRA datasets for large-scale reuse, with applications that extend beyond oncology transcriptomics. AVAILABILITY AND IMPLEMENTATION: Metappuccino source code is available on: github.com/chumphati/Metappuccino. The fine-tuned LLM, MetappuccinoLLModel, is available on: huggingface.co/chumphati/MetappuccinoLLModel. Both repositories are released under Apache-2.0 license.

Indexed as

Computational BiologyMetadataNeoplasmsSoftwareHigh-Throughput Nucleotide SequencingHumansLarge Language ModelsSequence Analysis, RNA

Identifiers

PMID42057294
PMCPMC13148957

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.