Evidence map›Paper›PMID 42539124›Full record

ArticleArXiv2026

Omics Data Discovery Agents: Agent-Supported Retrieval, Reanalysis, and Synthesis of Published Omics Data.

Alexandre Hutton, Jesse G Meyer

Abstract readPreprint
In one paragraph

Article in ArXiv, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Alexandre HuttonDepartment of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA, USA.
Jesse G MeyerDepartment of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA, USA.

Funding

Democratizing Multi-Omics to Expedite Discovery of Hidden Metabolic PathwaysR35GM142502 · NIGMS · MEDICAL COLLEGE OF WISCONSIN · PI MEYER, JESSE · 2021 to 2025
$2.2M
NIGMS NIH HHS R35 GM142502
6 · The paper itself

Abstract

The biomedical literature contains a vast collection of omics studies, yet most published data remain functionally inaccessible for computational reuse. When raw data are deposited in public repositories, essential information for reproducing reported results is dispersed across main text, supplementary files, and code repositories, and in the rarer cases where intermediate data (e.g. protein abundance files) are shared, their location is irregular. Here we present an agentic framework for the agent-supported retrieval, reanalysis, and synthesis of published omics data. The system employs large language model (LLM) agents with access to tools for fetching omics studies, extracting article metadata, identifying and downloading published data, executing containerized quantification pipelines, and synthesizing results across studies. Applied at corpus scale, the pipeline catalogued dataset references across thousands of PubMed Central articles; we report these as descriptive system outputs rather than as a validated measure of extraction accuracy. Using model context protocol (MCP) servers to expose containerized analysis tools, the agents retrieved and re-quantified data in five end-to-end reanalyses spanning data-dependent and data-independent proteomics and bulk RNA-seq. All five reanalyses completed, each with documented human guidance and workflow accommodations, and reproduced the authors' deposited abundances with high per-sample correlation (0.85-0.997) and strongly concordant differentially expressed features (fold-change Spearman 0.88-0.91), with no direction reversals among features called differentially expressed in both analyses; residual differences in significant-feature lists were attributable to threshold placement, tool-version, and preprocessing differences rather than to the underlying quantities. We further demonstrate that agents can identify semantically similar studies, judge data compatibility, and synthesize findings across studies, including a random-effects meta-analysis that recovered consistent protein regulation in liver fibrosis. Rather than a validated benchmark of literature-wide performance, this work is a feasibility demonstration together with an auditable, reusable toolset, establishing a foundation for prospective evaluation of automated omics-data reuse.

Indexed as

automated curationdata reanalysislarge language modelsmeta-analysismodel context protocolomics dataproteomicstranscriptomics

Identifiers

PMID42539124
PMCPMC13419620

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.