ArticleTomography (Ann Arbor, Mich.)2026
Temporal Reproducibility of Fracture Interpretation in Forensic Radiography: A Multispecialty Comparison of Physicians and Vision Language Models Including Fracture Subtype Description.
Article in Tomography (Ann Arbor, Mich.), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundAccurate fracture interpretation on plain radiographs is critical for both trauma care and medico-legal decision-making, where reproducibility is as important as point accuracy. Although vision language models (VLMs) have shown promising diagnostic performance, their temporal stability in forensic radiography remains unclear.
methodsWe analyzed 300 forensic radiographs (150 fracture-positive, 150 fracture-negative) from six long bones, independently evaluated by three emergency medicine physicians, three forensic medicine physicians, and three VLMs (ChatGPT-5.2, Gemini 3 Pro, Claude Sonnet 4.5) using an identical task format. Assessments included fracture presence and structured fracture subtype description (bone, morphology, displacement). VLM evaluations were repeated after one month under identical conditions.
resultsPhysician accuracy ranged from 79.7% to 98.7%. Emergency physicians reached the higher median sensitivity (92.0% against 76.7%), while specificity among the forensic readers was the more tightly clustered (median 92.0%, range 87.3-100.0%). ChatGPT-5.2 was the most accurate model (83.0%; sensitivity 70.0%, specificity 96.0%), followed by Gemini 3 Pro (77.7%), whereas Claude Sonnet 4.5 reached only 46.3% because of an extreme false-positive tendency (specificity 11.3%). Over one month, accuracy changed by -5.2, -4.1 and +8.6 percentage points, but these net figures concealed considerable case-level movement: within-model agreement ranged from near chance to substantial (mean Cohen κ 0.131 to 0.622), and the F1-score of Claude Sonnet 4.5 fell by 13.1 points despite its higher accuracy. Subtype descriptions were frequently correct once a fracture had been detected, but end-to-end subtype accuracy remained low.
conclusionsCurrent vision language models demonstrate encouraging diagnostic performance; however, their temporal reproducibility remains inadequate for independent medico-legal fracture interpretation. These findings highlight that reproducibility, in addition to diagnostic accuracy, should be considered a core benchmark when evaluating VLMs for high-stakes clinical and forensic use. Larger multicenter studies using independent external datasets are needed before forensic application is considered.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.