Evidence map›Paper›PMID 42440193›Full record

SynthesisNeurological sciences : official journal of the Italian Neurological Society and of the Italian Society of Clinical Neurophysiology2026

Generative large language models in the clinical management of Alzheimer's disease and mild cognitive impairment.

Yosef Adiniaev, Mahmud Omar, Oved Daniel, Tohar M Timor, Yiftach Barash, Olga R Brook, Eyal Klang, Alon Gorenshtein

Abstract readSystematic Review
In one paragraph

Synthesis in Neurological sciences : official journal of the Italian Neurological Society and of the Italian Society of Clinical Neurophysiology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Yosef AdiniaevFaculty of Medicine, University of Debrecen, Debrecen, Hungary. yosefad1305@gmail.com.ORCID http://orcid.org/0009-0002-7678-633X
Mahmud OmarBRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, 330 Brookline Ave, Boston, MA, 02115, USA.ORCID http://orcid.org/0009-0001-0438-0827
Oved DanielNeurology Division, Tel Aviv Sourasky University Medical Center, Tel Aviv, Israel.ORCID http://orcid.org/0000-0002-8232-9138
Tohar M TimorFaculty of Medicine, University of Debrecen, Debrecen, Hungary.ORCID http://orcid.org/0009-0000-9842-1636
Yiftach BarashBRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, 330 Brookline Ave, Boston, MA, 02115, USA.ORCID http://orcid.org/0000-0002-7242-1328
Olga R BrookDepartment of Radiology, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA.ORCID http://orcid.org/0000-0002-7074-8909
Eyal Klang *BRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, 330 Brookline Ave, Boston, MA, 02115, USA.ORCID http://orcid.org/0000-0002-4567-3108
Alon Gorenshtein *BRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, 330 Brookline Ave, Boston, MA, 02115, USA. agorensh@bidmc.harvard.edu.ORCID http://orcid.org/0009-0000-7542-8608

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundDementia affects over 55 million people worldwide. Mild cognitive impairment (MCI) often precedes Alzheimer's disease (AD). Clinical management requires integrating uncertain evidence from neuropsychological testing, neuroimaging, and biomarkers. Large language models (LLMs) also generate probabilistic outputs, but whether they can reliably support diagnostic, therapeutic, or educational tasks in AD and MCI has not been systematically examined.

methodsWe searched PubMed, Scopus, and PubMed Central (January 2023 to April 2026) for studies evaluating generative LLMs on clinical tasks in Alzheimer's disease (AD) or mild cognitive impairment (MCI). Risk of bias was assessed using QUADAS-AI and AXIS. Narrative synthesis followed the SWiM guideline. PROSPERO: CRD420261372436.

resultsEleven studies were included: diagnosis (n = 3), treatment guidance (n = 2), and patient/caregiver education (n = 8); two studies contributed to multiple domains. Diagnostic models achieved high internal accuracy (0.94-0.97) but declined on external validation; three-way classification accuracy dropped approximately 7% points, and MMSE-prediction R² collapsed from 0.90 to 0.25 on an external dataset. Treatment guidance approached but did not match structured clinical guidelines. Educational outputs were rated moderate to high quality but lacked source attribution and exceeded recommended reading levels; retrieval augmentation improved usability without improving accuracy. Hallucination was quantified in only 2 of 11 studies, and no study evaluated prospective clinical use.

conclusionsCurrent evidence does not support the use of LLMs for diagnosis, treatment selection, or patient education in AD/MCI without clinician oversight. These findings reflect the specific model versions, prompting strategies, and evaluation conditions in place at the time of each study, and are further limited by small heterogeneous evaluations, sparse hallucination measurement, and absence of prospective clinical validation.

Indexed as

Alzheimer DiseaseCognitive DysfunctionLarge Language ModelsGenerative Artificial IntelligenceHumansAlzheimer's diseaseClinical decision supportLarge language modelsMild cognitive impairmentPatient educationSystematic review

Identifiers

PMID42440193
PMCPMC13364785

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.