Evidence map›Paper›PMID 42620648›Full record

ArticleOphthalmology science2026

Transforming Systematic Reviews: Evaluating a Fine-Tuned Large Language Model for Abstract Screening in Uveitis and Retinal Vasculitis: Fine-Tuned LLM for Review Screening.

Carlos Cifuentes-González, Maxwell B Singer, William Rojas-Carabali, Germán Mejía-Salgado, Maria Vittoria Cicinelli, Jyotirmay Biswas, Sapna Gangaputra, Alejandra de-la-Torre, Vishali Gupta, Jose S Pulido and 1 more

Abstract read
In one paragraph

Article in Ophthalmology science, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Carlos Cifuentes-GonzálezNational Healthcare Group Eye Institute, Tan Tock Seng Hospital, Singapore, Singapore.
Maxwell B SingerDepartment of Ophthalmology and Visual Science, Yale School of Medicine, New Haven, Connecticut.
William Rojas-CarabaliNational Healthcare Group Eye Institute, Tan Tock Seng Hospital, Singapore, Singapore.
Germán Mejía-SalgadoNeuroscience Research Group (NEUROS), Neurovitae Center for Neuroscience, Institute of Translational Medicine (IMT), Escuela de Medicina y Ciencias de la Salud, Universidad del Rosario, Bogotá, Colombia.
Maria Vittoria CicinelliSchool of Medicine, Vita-Salute San Raffaele University, Milan, Italy.
Jyotirmay BiswasDepartment of Uveitis Services, Sankara Nethralaya, Chennai, India.
Sapna GangaputraVanderbilt Eye Institute, Vanderbilt University Medical Center, Nashville, Tennessee.
Alejandra de-la-TorreNeuroscience Research Group (NEUROS), Neurovitae Center for Neuroscience, Institute of Translational Medicine (IMT), Escuela de Medicina y Ciencias de la Salud, Universidad del Rosario, Bogotá, Colombia.
Vishali GuptaAdvanced Eye Centre, Postgraduate Institute of Medical Education and Research (PGIMER), Chandigarh, India.
Jose S PulidoRetina Service, Wills Eye Hospital, Thomas Jefferson University, Philadelphia, Pennsylvania.
Rupesh AgrawalNational Healthcare Group Eye Institute, Tan Tock Seng Hospital, Singapore, Singapore.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Purpose: To evaluate the classification performance of UveAItis, a domain-specific large language model (LLM) fine-tuned for automated title and abstract screening in systematic reviews, using retinal vasculitis as a prototype. Design: Comparative evaluation study embedded within a registered systematic review and meta-analysis (PROSPERO: CRD42023489232). Subjects: A total of 1030 randomly selected articles from an initial search of 5533 records related to retinal vasculitis. Methods: Articles were independently screened by 2 uveitis experts (gold standard), final-year medical students, and 3 LLMs: UveAItis (fine-tuned Generative Pre-trained Transformer [GPT]-4o), base GPT-4o, and Claude Sonnet 3.5. Screening followed a 2-question binary logic regarding human subjects and primary empirical research design. Discrepancies were resolved through expert adjudication. Main Outcome Measures: Classification accuracy, sensitivity, specificity, area under the receiver operating characteristic curve, and Cohen Kappa coefficient for inter-rater agreement. Results: UveAItis achieved the highest performance with an accuracy of 93.3%, area under the curve (AUC) of 0.887, and Kappa of 0.77. It significantly outperformed base GPT-4o (AUC: 0.805, Conclusions: UveAItis demonstrated expert-level performance, significantly outperforming general-purpose LLMs and nonexpert human reviewers. These findings validate the potential of domain-specific fine-tuning to enhance the efficiency, scalability, and reproducibility of evidence synthesis in specialized medical fields like ophthalmology. Financial Disclosures: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

Indexed as

Artificial intelligenceLarge language modelRetinal vasculitisSystematic reviewUveitis

Identifiers

PMID42620648
PMCPMC13485612

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.