Evidence map›Paper›PMID 42758740›Full record

ArticlePLOS digital health2026

The economics of accuracy for medical reasoning with large language models.

Kiran Bhattacharyya, Sreeram Kamabattula

Abstract read
In one paragraph

Article in PLOS digital health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Kiran BhattacharyyaSKA Labs, Atlanta, Georgia, United States of America.ORCID https://orcid.org/0000-0002-1776-4782
Sreeram KamabattulaAdvanced Product Development, Intuitive Surgical, Inc., Atlanta, Georgia, United States of America.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Deploying large language models (LLMs) in clinical settings is limited by security, reliability, latency, and accessibility concerns that favor smaller, on-device or on-premise models. However, these smaller models may struggle to meet accuracy requirements. While fine-tuning and retrieval-augmented generation (RAG) can improve domain-specific accuracy, these methods require additional labeled data, technical skill, and infrastructure. In contrast, test-time scaling-allocating extra token-budget during inference-offers a training-free alternative to increasing accuracy. However, the trade-offs between these strategies and their interaction with model size remain poorly understood for medical reasoning. To address this gap, we compare three approaches-test-time scaling, fine-tuning, and context grounding-using the Gemma and MedGemma family of LLMs (Gemma-3 1B, Gemma-3 4B, Gemma-3 27B, MedGemma-4B, and MedGemma 27B) and evaluate these systems across common biomedical question-answering (QA) datasets and a set of recently released medical exam questions with the performance of practicing clinicians available for comparison. We test baseline prompts (direct answer, Chain-of-Thought, and self-consistency) while introducing a new prompting method we call "prompt-chaining for continuous reflection" (PCCR) that forces inference time minimum token-generation budgets. We assess accuracy and tokens-generated, allowing us to investigate the accuracy-efficiency trade-offs across prompting, context-grounding, fine-tuning, and model scales. We discover equivalency point configurations where a smaller model's accuracy falls within one 95% confidence interval of a larger model's (typically within 1-4 percentage points) reached through increased reasoning budgets, context-grounding, or fine-tuning. Specific effects are apparent and statistically supported: the benefit of medical fine-tuning grew from +4.6 to +15.7 percentage points (non-overlapping 95% CIs) when paired with self-consistency, and enforced extended reasoning raised MedGemma 27B accuracy from 58.1% to 80.1% (p < 10-5) in the absence of context. We also identify an "overthinking" inflection: when high-quality context is available, extended reasoning beyond roughly 128-256 tokens degrades accuracy by 7-13 percentage points. Using these empirical results, we formulate a general framework with equations to balance cost-benefit trade-offs when engineering LLM-based systems for medical reasoning and QA. We recommend generalizable configurations, designs, and patterns to achieve accuracy and efficiency objectives for example use-cases relevant to healthcare organizations.

Identifiers

PMID42758740
PMCPMC13588378

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.