ArticleScientific reports2026
Benchmarking large language models on persian surgical subspecialty board examinations: a comparative study of ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash.
Shahab Sheikhalishahi et al.PubMed ↗Full text ↗Publisher ↗
No numbers read from the abstract.