Evidence map›Paper›PMID 42801061›Full record

ArticleJAMIA open2026

Large language model-based triage to identify antiretroviral therapy adherence barriers and risk levels in patient messages.

Yuanchao Ma, Sofiane Achiche, David Lessard, Kim Engler, Serge Vicente, Gavin Tu, Benoît Lemire, Lina Del Balso, Nathalie Paisible, MARVIN Chatbots Patient Expert Committee and 1 more

Abstract read
In one paragraph

Article in JAMIA open, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Yuanchao MaInstitute of Biomedical Engineering, Polytechnique Montreal, Montreal, QC H3T 1J4, Canada.ORCID https://orcid.org/0000-0002-4048-1705
Sofiane AchicheInstitute of Biomedical Engineering, Polytechnique Montreal, Montreal, QC H3T 1J4, Canada.
David LessardCentre for Outcomes Research & Evaluation, Research Institute of the McGill University Health Centre, Montreal, QC H4A 3J1, Canada.
Kim EnglerCentre for Outcomes Research & Evaluation, Research Institute of the McGill University Health Centre, Montreal, QC H4A 3J1, Canada.
Serge VicenteDépartement des enseignements généraux, École de technologie supérieure, Université du Québec, Montreal, QC H3C 1K3, Canada.
Gavin TuFaculty of Medicine, Université Laval, Quebec, QC G1V 0A6, Canada.
Benoît LemireChronic Viral Illness Service, Division of Infectious Disease, Department of Medicine, McGill University Health Centre, Montreal, QC H4A 3J1, Canada.
Lina Del BalsoChronic Viral Illness Service, Division of Infectious Disease, Department of Medicine, McGill University Health Centre, Montreal, QC H4A 3J1, Canada.
Nathalie PaisibleChronic Viral Illness Service, Division of Infectious Disease, Department of Medicine, McGill University Health Centre, Montreal, QC H4A 3J1, Canada.
MARVIN Chatbots Patient Expert Committee
Bertrand LebouchéCentre for Outcomes Research & Evaluation, Research Institute of the McGill University Health Centre, Montreal, QC H4A 3J1, Canada.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objective: This study aimed to develop large language models (LLMs) to automatically identify antiretroviral therapy (ART) adherence barriers and stratify nonadherence risk levels from patient-generated messages. Materials and Methods: With a co-construction committee of people with HIV and providers, 15 480 sentences were annotated for barrier and risk levels. General-domain LLMs (eg, Flan-T5) and clinical foundation models (eg, Clinical-T5) were fine-tuned for multiclass classification and evaluated using Macro-F1. Best-performing models were compared with GPT family and other open-source LLMs. Model fairness, error patterns, and environmental footprints were also assessed. Results: Flan-T5-xl achieved the best barrier detection (Macro-F1 = 0.83 test/0.71 external), and Flan-T5-large excelled in risk stratification (0.79/0.57). Fine-tuned general-domain LLMs significantly outperformed clinical foundation models ( Discussion: Fine-tuned Flan-T5 models demonstrated strong classification performance, greater robustness to demographic attributes, and lower energy consumption, though challenges remained for subjective and underrepresented categories, reflecting both data imbalance and model limitations in implicit reasoning. Conclusion: LLM-based approaches show promise for real-time ART adherence monitoring, offering a scalable solution to individualized HIV care. Beyond performance, our findings highlight the importance of fairness and environmental sustainability in clinical AI development, with next steps focused on real-world validation through deployment in patient-facing digital tools.

Indexed as

clinical triageHIVlarge language modelsmedication adherencenatural language processingresponsible AI

Identifiers

PMID42801061
PMCPMC13615712

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.