Evidence map›Paper›PMID 41584932›Full record

ArticleHealth information science and systems2026

Large language models and conditional rules in clinical decision support systems.

Shangeetha Sivasothy, Adrian Bingham, Irini Logothetis, Scott Barnett, Mohamed Abdelrazek, Carl Luckhoff, Joseph Mathew, Rajesh Vasa, Kon Mouzakis

Abstract read
In one paragraph

Article in Health information science and systems, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Shangeetha SivasothyApplied Artificial Intelligence Initiative, Deakin University, Geelong, VIC Australia.
Adrian BinghamApplied Artificial Intelligence Initiative, Deakin University, Geelong, VIC Australia.
Irini LogothetisApplied Artificial Intelligence Initiative, Deakin University, Geelong, VIC Australia.ORCID 0000-0003-0143-3812
Scott BarnettApplied Artificial Intelligence Initiative, Deakin University, Geelong, VIC Australia.
Mohamed AbdelrazekApplied Artificial Intelligence Initiative, Deakin University, Geelong, VIC Australia.
Carl LuckhoffEmergency and Trauma Centre, Alfred Health, Melbourne, VIC Australia.
Joseph MathewEmergency and Trauma Centre, Alfred Health, Melbourne, VIC Australia.
Rajesh VasaApplied Artificial Intelligence Initiative, Deakin University, Geelong, VIC Australia.
Kon MouzakisApplied Artificial Intelligence Initiative, Deakin University, Geelong, VIC Australia.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Clinical Decision Support Systems (CDSS) improve patient outcomes and support sustainable health services by enhancing medical decisions. Developing rules for a CDSS is expensive due to delays in capturing and defining the rules through multiple iterations between clinicians and developers as the role of a clinician is patient care. Objective: We investigate the effectiveness of large language models (LLMs) and large reasoning models (LRMs) in generating a triaging rule set for a CDSS. Methods: We prompt various LLMs (GPT-3.5, GPT-4, GPT-4o, Gemini, Claude 3.5 Sonnet) and various LRMs (GPT-o1-mini, Grok-4, Claude 4 Sonnet) using alternative prompting techniques. We compare the LLM generated rule sets against the clinical rule set from our Pandemic Intervention Monitoring System (PiMS); a triaging CDSS built in collaboration with clinicians to monitor COVID-19 positive patients. Effectiveness is evaluated based on the accuracy, interpretability, and rule complexity. Results: We identified that LLMs generated COVID-19 screening rule sets compared to triaging rule sets when not specifying the variables from our PiMS rule set. By including PiMS variables in our prompts, we discovered LLMs 1) had lower interpretability and rule complexity compared to the PiMS rule set, and 2) resulted in an average accuracy between 31.62% ± 0.19% and 70.71% ± 0.02%. While for LRMs, we identified that 1) interpretability varied between 3 and 94 compared to 41 identified in our PiMS rule set and 2) resulted in an average accuracy between 31.62% ± 0.19% and 81.70 ± 0.05%. Conclusions: LLMs are limited in emulating clinical rule sets due to their simplicity and lack of complex reasoning. Despite LRMs improving effectiveness, they are still limited. LLMs and LRMs can generate a feasible initial rule set for CDSS. This can reduce time invested by clinicians and developers by minimising the number of iterations for refinement. Future work can explore integrating LLMs and LRMs with decision trees to improve effectiveness. Supplementary Information: The online version contains supplementary material available at 10.1007/s13755-026-00428-z.

Indexed as

Clinical decision support systemsConditional rulesLarge language models

Identifiers

PMID41584932
PMCPMC12824032

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.