Evidence map›Paper›PMID 42415081›Full record

ArticleBMC medical informatics and decision making2026

Generating Alzheimer's narratives using large language models.

Paula Andrea Perez-Toro, Mahmoud Almizel, Elmar Nöth, Andreas Maier, Tomas Arias-Vergara

Abstract read
In one paragraph

Article in BMC medical informatics and decision making, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Paula Andrea Perez-ToroPattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany. paula.andrea.perez@fau.de.ORCID 0000-0002-2727-2116
Mahmoud AlmizelPattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany.
Elmar NöthPattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany.
Andreas MaierPattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany.
Tomas Arias-VergaraPattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundAnalyzing semi-spontaneous speech is a promising direction for supporting Alzheimer's disease (AD) assessment, yet progress is limited by the scarcity of annotated clinical data. Large Language Models (LLMs) offer new opportunities to generate synthetic narratives that may resemble speech patterns of both patients with AD and healthy controls during cognitive evaluation tasks such as the Cookie Theft Picture description.

methodsThis study evaluates whether models including GPT, T5/Flan-T5, LLaMA, Mistral, and Qwen can generate clinically plausible picture-description narratives under two configurations: Human-to-Bot, where an LLM responds directly to real interviewer prompts, and Bot-to-Bot, where two LLMs simulate both interviewer and participant roles. Models were fine-tuned on transcripts from the DementiaBank Pitt Corpus and assessed using lexical and semantic metrics, as well as human expert ratings. Generated narratives were further used to augment training data for an AD vs. healthy control classifier based on BERT embeddings and an MLP architecture.

resultsLLMs differed substantially in their ability to reproduce clinically meaningful and semantically coherent narratives of patient-interviewer interactions. Mistral, LLaMA, and Qwen achieved the strongest automatic evaluation metrics, e.g., BERTScores above 0.90 in the Human-to-Bot condition-and produced narratives rated by human experts as fluent, plausible, and diagnostically informative. When combining real and synthetic narratives for classifier training, the highest F1-score reached 0.84, outperforming models trained on real data alone (F1 = 0.74). Synthetic data generated in Human-to-Bot settings contributed most to diagnostic improvements, whereas Bot-to-Bot interactions exhibited greater variability and reduced clinical realism.

conclusionLLMs can generate high-quality synthetic narratives that enhance downstream AD classification and show promising clinical plausibility in cognitive assessment contexts. Incorporating LLM-generated data provides a scalable strategy for mitigating data scarcity in dementia research. Future work should focus on improving fully synthetic dialogue quality, expanding multilingual capabilities, and refining evaluation frameworks to better capture clinically relevant linguistic features.

Indexed as

Alzheimer DiseaseLarge Language ModelsNarrationHumansAlzheimer’s diseaseGenerative AILarge language modelsSynthetic data augmentation

Identifiers

PMID42415081
PMCPMC13352925

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.