Evidence map›Paper›PMID 42684322›Full record

ArticleJMIR formative research2026

Toward Automating the Selection of Articles Reporting EQ-5D Data for Systematic Literature Reviews Using Large Language Models: Algorithm Development and Evaluation Study.

Gábor Kertész, János Tibor Czere, Zsombor Zrubka, Laszlo Gulacsi, Marta Pentek

Abstract read
In one paragraph

Article in JMIR formative research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Gábor KertészSoftware Engineering Institute, John von Nemann Faculty of Informatics, Obuda University, Becsi str 96/B, Budapest, 1034, Hungary, +3616665528.ORCID 0000-0002-8845-8301
János Tibor CzereDoctoral School of Innovation Management, Obuda University, Budapest, Hungary.ORCID 0000-0002-9432-7906
Zsombor ZrubkaDoctoral School of Innovation Management, Obuda University, Budapest, Hungary.ORCID 0000-0002-1992-6087
Laszlo GulacsiDoctoral School of Innovation Management, Obuda University, Budapest, Hungary.ORCID 0000-0002-9285-8746
Marta PentekDoctoral School of Innovation Management, Obuda University, Budapest, Hungary.ORCID 0000-0001-9636-6012

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Systematic literature reviews (SLRs) are essential for evidence synthesis in health research but remain labor-intensive, especially at the screening stage. Manual review of titles and abstracts requires substantial human effort, while existing automation tools still have limited adoption in health technology assessment. The EQ-5D questionnaire, a widely used patient-reported outcome measure for health-related quality of life, provides data that frequently underpin reimbursement and policy decisions. Objective: This pilot study evaluated whether recent large language models (LLMs) can support the identification of publications reporting EQ-5D data in PubMed records, using only publicly available metadata (title, abstract, and keywords). Methods: A total of 200 publications retrieved through the EuroQol PubMed filter were manually labeled by experts as reporting or not reporting EQ-5D data. The dataset was split into stratified training, validation, and test subsets. Several machine learning approaches were compared, including a Naïve Bayes baseline using bag-of-words features, a decision-tree model based on full-text keyword occurrence, and transformer-based LLMs (Bidirectional Encoder Representations from Transformers [BERT], Biomedical BERT [BioBERT], Scientific BERT [SciBERT], and Biomedical Language Understanding Evaluation BERT [BlueBERT]). Both classifier-only and fine-tuned configurations were tested across multiple learning rates. Model performance was assessed using accuracy, precision, recall, and Results: Baseline approaches achieved near-random test performance (accuracy around 0.53). Classifier-only LLMs modestly improved results (accuracy up to 0.64 with SciBERT). Fine-tuned models substantially outperformed these baselines, with BERT and BioBERT achieving the best performance (accuracy=0.70; Conclusions: This study provides the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature. The findings support technical feasibility but do not establish a reliable stand-alone automated screening tool. Although limited by dataset size, the proposed workflow is reproducible and adaptable to other patient-reported outcome measures. Because validation was based on a single small train-validation-test split, the results should be interpreted as preliminary; future work will scale data collection, include statistical testing, and explore semisupervised learning to further reduce manual screening workload.

Indexed as

AlgorithmsLarge Language ModelsSystematic Reviews as TopicHumansPilot ProjectsSurveys and QuestionnairesBERTBidirectional Encoder Representations from TransformersBioBERTBiomedical Bidirectional Encoder Representations from TransformersEQ-5Dhealth informaticslarge language modelPubMedquality of lifesystematic literature review

Identifiers

PMID42684322
PMCPMC13528701

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.