Evidence map›Paper›PMID 42365257›Full record

ArticleBMC oral health2026

Prompt sensitivity of large language models in orthodontic patient counseling: a scenario-based experimental study.

Ersin Yıldırım, Esra Tunalı, Şeniz Karaçay

Abstract read
In one paragraph

Article in BMC oral health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Ersin YıldırımDepartment of Orthodontics, Hamidiye Faculty of Dentistry, University of Health Sciences, Istanbul, Türkiye. ersin.yildirim@sbu.edu.tr.ORCID 0000-0003-0142-246X
Esra TunalıDepartment of Orthodontics, Hamidiye Faculty of Dentistry, University of Health Sciences, Istanbul, Türkiye.
Şeniz KaraçayDepartment of Orthodontics, Hamidiye Faculty of Dentistry, University of Health Sciences, Istanbul, Türkiye.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundLarge language models (LLMs) are increasingly used as accessible sources of health information, including orthodontic patient counseling. While previous studies have evaluated the accuracy and reliability of AI-generated responses, the effect of prompt formulation on the safety and quality of orthodontic advice remains unclear. Understanding prompt sensitivity is essential for assessing the real-world reliability of conversational AI systems, as patient queries are typically expressed in diverse linguistic forms.

methodsThis in silico experimental study evaluated prompt sensitivity using 24 standardized orthodontic clinical scenarios. Each scenario was queried using four prompt formulations (brief layperson, detailed patient, professional clinical, and anxiety-driven), resulting in 96 prompts. These were submitted to four LLMs (ChatGPT, Gemini, Copilot, and Claude), generating 384 responses. Responses were independently evaluated by two orthodontic experts using a predefined expert scoring rubric across accuracy, safety, completeness, and clarity. Consensus scores were analyzed using the Friedman test for prompt effects and the Kruskal-Wallis test for model comparisons. Unsafe response rates and prompt robustness indices were also calculated.

resultsSafety scores did not differ significantly across prompt formulations (χ²(3) = 3.40, p = 0.334) or between models (H = 0.17, p = 0.982). A total of 5 of 384 responses (1.3%) were classified as unsafe. Prompt robustness analysis demonstrated low variability (mean prompt robustness index = 0.15). Response length differed significantly across prompt types (p = 0.0029), whereas response time did not (p = 0.998). A significant difference in clarity scores was observed across models (p = 0.029), with post hoc analysis indicating higher clarity scores for ChatGPT than Claude (adjusted p = 0.020).

conclusionsLLMs demonstrated consistent and clinically safe performance in orthodontic patient counseling, with minimal sensitivity to prompt formulation. While prompt wording influenced response length, it did not affect clinical reliability. Differences between models were primarily related to clarity rather than content. LLMs may provide stable informational support across diverse patient queries; however, their outputs should remain adjunctive to professional care.

Indexed as

CounselingLarge Language ModelsOrthodonticsHumansReproducibility of ResultsArtificial intelligenceChatGPTDigital healthLarge language modelsOrthodonticsPatient educationPrompt sensitivity

Identifiers

PMID42365257
PMCPMC13508204

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.