Evidence map›Paper›PMID 42592216›Full record

ArticleFrontiers in digital health2026

Performance and limitations of four large language models in genetic counseling for thalassemia.

Wenfu Zhong, Jingwen Huang, Mengsi Wei, Qingpeng Liang, Yuanwu Yang, Jinrong Chen, Jinjiang Mao, Ying Qin, Yifan Sun, Yishan Liang

Abstract read
In one paragraph

Article in Frontiers in digital health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Wenfu Zhong *Department of Clinical Laboratory, Guigang Clinical Medical Research Center for Medical Laboratory, Guigang City People's Hospital, Guigang, Guangxi, China.
Jingwen Huang *Department of Respiratory and Critical Care Medicine, Guigang City People's Hospital, Guigang, Guangxi, China.
Mengsi WeiDepartment of Clinical Laboratory, Guigang Clinical Medical Research Center for Medical Laboratory, Guigang City People's Hospital, Guigang, Guangxi, China.
Qingpeng LiangDepartment of Clinical Laboratory, Guigang Clinical Medical Research Center for Medical Laboratory, Guigang City People's Hospital, Guigang, Guangxi, China.
Yuanwu YangDepartment of Clinical Laboratory, Guigang Clinical Medical Research Center for Medical Laboratory, Guigang City People's Hospital, Guigang, Guangxi, China.
Jinrong ChenDepartment of Clinical Laboratory, Guigang Clinical Medical Research Center for Medical Laboratory, Guigang City People's Hospital, Guigang, Guangxi, China.
Jinjiang MaoDepartment of Obstetrics, Guigang City People's Hospital, Guigang, Guangxi, China.
Ying QinDepartment of Obstetrics, Guigang City People's Hospital, Guigang, Guangxi, China.
Yifan SunDepartment of Clinical Laboratory, Guigang Clinical Medical Research Center for Medical Laboratory, Guigang City People's Hospital, Guigang, Guangxi, China.
Yishan LiangDepartment of Clinical Laboratory, Guigang Clinical Medical Research Center for Medical Laboratory, Guigang City People's Hospital, Guigang, Guangxi, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

In regions with high thalassemia prevalence, such as southern China and Southeast Asia, chronic shortages of professional genetic counseling resources have driven interest in large language models (LLMs) as auxiliary tools, yet their performance and safety boundaries in this setting remain uncharacterized. This single-center retrospective study evaluated four LLMs (ChatGPT-5.2 Thinking, DeepSeek-V3.2 chat, Gemini 3 Flash, and Grok 4.1 Fast) using 1,080 standardized knowledge questions administered across five independent sessions and 150 real-world clinical cases scored by six senior experts across eight dimensions. All models exceeded 90% accuracy on single-choice and true-false questions. Between-model differences were most pronounced in multiple-choice questions, where ChatGPT-5.2 Thinking achieved the highest accuracy (87.28% ± 1.69%), significantly outperforming Grok 4.1 Fast (72.06% ± 1.69%,

Indexed as

genetic counselinglarge language modelsperformance evaluationprenatal diagnosisthalassemia

Identifiers

PMID42592216
PMCPMC13464380

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.