Evidence map›Paper›PMID 41959814›Full record

ArticlemedRxiv : the preprint server for health sciences2026

Combining Token Classification With Large Language Model Revision for Age-Friendly 4M Entity Recognition From Nursing Home Text Messages: Development and Evaluation Study.

Philip Amewudah, Mihail Popescu, Matthew S Farmer, Kimberly R Powell

Abstract readPreprint
In one paragraph

Article in medRxiv : the preprint server for health sciences, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Philip AmewudahInstitute for Data Science and Informatics, University of Missouri, Columbia, Columbia, USA.ORCID 0000-0003-3411-6143
Mihail PopescuDepartment of Biomedical Informatics, Biostatistics and Medical Epidemiology, University of Missouri School of Medicine, Columbia, USA.ORCID 0000-0002-6145-8096
Matthew S FarmerSchool of Nursing, University of Missouri, Columbia, Columbia, USA.ORCID 0000-0003-0989-2968
Kimberly R PowellSinclair School of Nursing, University of Missouri, Columbia, Missouri, USA.ORCID 0000-0002-6144-1438

Funding

Reducing Avoidable Nursing Home-to-Hospital Transfers of Residents with ADRD: An Analysis of Interdisciplinary Team Communication using Text MessagesR01AG078281 · NIA · UNIVERSITY OF MISSOURI-COLUMBIA · PI POWELL, KIMBERLY RYAN · 2022 to 2024
$1.1M
NIA NIH HHS R01 AG078281
6 · The paper itself

Abstract

Background: Secure text messages (TMs) exchanged among interdisciplinary care teams in nursing homes (NHs) contain clinical information that aligns with the Age-Friendly Health Systems 4Ms: What Objective: This study aimed to develop and evaluate a multi-stage 4M Entity Recognition (4M-ER) pipeline that combines a fine-tuned token classifier with large language model (LLM) revision, using only locally deployed open-source models, to improve 4M information extraction from clinical TMs. Methods: We used an expert-annotated dataset of 1,169 TMs representing conversations between interdisciplinary care teams across 16 Midwest NHs. The pipeline first identifies candidate text spans using a fine-tuned Bio-ClinicalBERT token classifier. A semantic similarity retriever then selects in-context exemplars to guide an LLM revision in which the LLM (Gemma, Phi, Qwen, or Mistral) performs boundary correction, label evaluation, and selective acceptance or rejection of candidate spans. Baselines for comparison included single-stage zero-shot LLMs, single-stage fine-tuned Bio-ClinicalBERT, and a fine-tuned LLM (Gemma) from a prior study. Ablation studies assessed the contribution of each pipeline stage and the effect of message filtering. Robustness was evaluated across 5 repeated runs. Results: The 4M-ER pipeline outperformed the previously fine-tuned Gemma LLM across all 4M domains, achieving F Conclusions: The 4M-ER pipeline enables accurate and scalable extraction of 4M entities from clinical TMs by combining fine-tuned Bio-ClinicalBERT with LLM revision using only locally deployed open-source models. The structured 4M data produced by the pipeline can support 4M taxonomy and ontology construction, as demonstrated in the prior work, and provides a foundation for downstream applications including real-time clinical surveillance, compliance with emerging age-friendly quality measures, and predictive modeling in long-term care settings.

Indexed as

Age-Friendly Health SystemsInformation ExtractionLarge Language ModelsLong-Term CareNamed Entity RecognitionNatural Language ProcessingNursing HomesNursing InformaticsText Messaging

Identifiers

PMID41959814
PMCPMC13060489

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.