Evidence map›Paper›PMID 41886735›Full record

ArticleJMIR AI2026

Large Language Model Adaptation Strategies in Speech-Based Cognitive Screening: Systematic Evaluation.

Fatemeh Taherinezhad, Mohamad Javad Momeni Nezhad, Sepehr Karimi, Sina Rashidi, Ali Zolnour, Maryam Dadkhah, Yasaman Haghbin, Hossein Azadmaleki, Maryam Zolnoori

Abstract read
In one paragraph

Article in JMIR AI, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Fatemeh Taherinezhad *Columbia University Irving Medical Center, 622 W, 168th St, New York, NY, 10032, United States.ORCID http://orcid.org/0009-0004-5650-6374
Mohamad Javad Momeni Nezhad *Columbia University Irving Medical Center, 622 W, 168th St, New York, NY, 10032, United States.ORCID http://orcid.org/0009-0001-9662-6133
Sepehr KarimiColumbia University Irving Medical Center, 622 W, 168th St, New York, NY, 10032, United States.ORCID http://orcid.org/0009-0001-8078-8955
Sina RashidiColumbia University Irving Medical Center, 622 W, 168th St, New York, NY, 10032, United States.ORCID http://orcid.org/0009-0009-1937-3396
Ali ZolnourColumbia University Irving Medical Center, 622 W, 168th St, New York, NY, 10032, United States.ORCID http://orcid.org/0009-0007-0668-5169
Maryam DadkhahColumbia University Irving Medical Center, 622 W, 168th St, New York, NY, 10032, United States.ORCID http://orcid.org/0009-0003-9769-5979
Yasaman HaghbinColumbia University Irving Medical Center, 622 W, 168th St, New York, NY, 10032, United States.ORCID http://orcid.org/0009-0005-0160-6113
Hossein AzadmalekiColumbia University Irving Medical Center, 622 W, 168th St, New York, NY, 10032, United States.ORCID http://orcid.org/0009-0008-4284-9869
Maryam ZolnooriColumbia University Irving Medical Center, 622 W, 168th St, New York, NY, 10032, United States.ORCID http://orcid.org/0000-0003-4484-2990

Funding

Development of a Screening Algorithm for Timely Identification of Patients with Mild Cognitive Impairment and Early Dementia in Home HealthcareR00AG076808 · NIA · COLUMBIA UNIVERSITY HEALTH SCIENCES · PI Maryam Zolnoori · 2024 to 2026
$742k
NIA NIH HHS R00 AG076808
6 · The paper itself

Abstract

Background: Over half of US adults with Alzheimer disease and related dementias (ADRD) remain undiagnosed. Speech-based screening algorithms offer a scalable approach, but the relative value of large language model (LLM) adaptation strategies is unclear. Objective: The study aimed to compare LLM adaptation strategies for cognitive impairment detection across DementiaBank speech datasets using both text-only and multimodal models. Methods: We analyzed audio-recorded speech from 237 participants in the ADReSSo subset of DementiaBank (ADRD vs cognitive normal [CN]) and report performance on a held-out test set (n=71). Nine text-only LLMs (3B-405B; open-weight and commercial) and 3 multimodal audio-text models were evaluated. Adaptations included (1) in-context learning (ICL) with 4 demonstration selection strategies (most similar, least similar, average similar or prototype, and random), (2) reasoning-augmented prompting (self- or teacher-generated rationales, self-consistency, tree-of-thought with domain experts), (3) parameter-efficient fine-tuning (token-level vs added classification head), and (4) multimodal audio-text integration. Generalizability of the adaptation strategies was evaluated on the DementiaBank Delaware dataset (n=205; mild cognitive impairment vs CN) using the first 3 strategies. The primary outcome was the F1-score for the cognitive impaired class; the area under the receiver operating characteristic curve was reported when available. Results: On the ADReSSo dataset, average similar (prototype) demonstrations achieved the highest ICL performance across model sizes (F1-score up to 0.81). Reasoning primarily benefited smaller models: teacher-generated rationales increased LLaMA 8B from F1-score 0.72 to 0.76; expert-role tree-of-thought improved its zero-shot score from 0.65 to 0.71. Token-level fine-tuning produced the highest scores (LLaMA 3B: F1=0.83, 95% CI 0.01, area under the curve [AUC]=0.91; LLaMA 70B: F1=0.82, 95% CI 0.02, AUC=0.86; GPT-4o: F1=0.79, 95% CI 0.01, AUC=0.87). A classification head markedly improved MedAlpaca 7B (F1=0.06, 95% CI 0.02 to F1=0.81, 95% CI 0.04), indicating model-dependent benefits of this approach. Among multimodal models, fine-tuned Phi-4 Multimodal reached an F1-score of 0.80 (cognitive impaired) and 0.75 (CN) but did not exceed the top text-only systems. On the Delaware dataset, ICL achieved a high performance (LLaMA 8B: F1=0.74; GPT-4o: F1=0.80). Reasoning-augmented ICL improved LLaMA 8B to an F1-score of 0.75. Token-level fine-tuning produced the highest scores (LLaMA 8B: F1=0.76, 95% CI 0.02; GPT-4o: F1=0.82, 95% CI 0.03). Conclusions: Detection accuracy is influenced by demonstration selection, reasoning design, and tuning method. Token-level fine-tuning is generally most effective, while a classification head benefits models that perform poorly under token-based supervision. Properly adapted open-weight models can match or exceed commercial LLMs, supporting their use in scalable speech-based ADRD and mild cognitive impairment screening. Current multimodal models may require improved audio-text alignment and/or larger training corpora.

Indexed as

cognitive impairment detectionfine-tuningin-context learninglarge language models adaptationmultimodal speech-text analysisreasoning-augmented promptingspeech-based screening

Identifiers

PMID41886735
PMCPMC13021110

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.