Evidence map›Paper›PMID 42715407›Full record

ArticleJournal of medical Internet research2026

Large Language Models for Distress Rating in Korean Psycho-Oncology Interviews: Exploratory Clinician-Benchmarked Evaluation Study.

Jaehyun Kim, Kyung-Lak Son, Chan-Woo Yeom, Won-Hyoung Kim, Sun Hyung Lee, Joon Sung Shin, Hyunsun Yang, Daehun Yoo, Bong-Jin Hahm

Abstract read
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Jaehyun KimDepartment of Neuropsychiatry, Seoul National University Hospital, 101, Daehak-ro, Jongno-gu, Seoul, 03080, Republic of Korea, 82 2-2072-2557.ORCID http://orcid.org/0000-0002-6553-9563
Kyung-Lak SonSeoul Rhema Psychiatric Clinic, Gangnam-gu, Seoul, Republic of Korea.ORCID http://orcid.org/0000-0002-9332-8659
Chan-Woo YeomDepartment of Psychiatry, Eulji University Uijeongbu Eulji Medical Center, Uijeongbu, Gyeonggi-do, Republic of Korea.ORCID http://orcid.org/0000-0003-2265-3291
Won-Hyoung KimDepartment of Psychiatry, Inha University Hospital, Jung-gu, Incheon, Republic of Korea.ORCID http://orcid.org/0000-0002-6650-3685
Sun Hyung LeeDepartment of Neuropsychiatry, Seoul National University Hospital, 101, Daehak-ro, Jongno-gu, Seoul, 03080, Republic of Korea, 82 2-2072-2557.ORCID http://orcid.org/0000-0001-8559-8117
Joon Sung ShinDepartment of Neuropsychiatry, Seoul National University Hospital, 101, Daehak-ro, Jongno-gu, Seoul, 03080, Republic of Korea, 82 2-2072-2557.ORCID http://orcid.org/0000-0002-2425-4406
Hyunsun YangGenesis Lab Inc, Jung-gu, Seoul, Republic of Korea.ORCID http://orcid.org/0009-0003-2578-7567
Daehun YooGenesis Lab Inc, Jung-gu, Seoul, Republic of Korea.ORCID http://orcid.org/0009-0000-7177-3885
Bong-Jin HahmDepartment of Neuropsychiatry, Seoul National University Hospital, 101, Daehak-ro, Jongno-gu, Seoul, 03080, Republic of Korea, 82 2-2072-2557.ORCID http://orcid.org/0000-0002-2366-3275

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Psychiatric distress is common among patients with cancer; yet, systematic interview-based screening remains difficult to scale in routine clinical care. Large language models (LLMs) have shown promise as scalable tools for mental health assessment, but most existing evidence is derived from clinician-authored records, translated text, or proxy data. The performance characteristics, error patterns, and explanatory behaviors of contemporary LLMs when applied to authentic, non-English psychiatric interviews remain insufficiently characterized. Objective: This exploratory study evaluated how contemporary LLMs reproduced psycho-oncologists' item-level symptom ratings from real-world Korean psycho-oncology interviews, focusing on concordance, directional bias, and clinician-adjudicated error characteristics. Methods: Between April 2024 and May 2025, 101 adults receiving oncologic care in South Korea underwent semistructured interviews. Board-certified psycho-oncologists provided real-time, time-stamped ratings of 39 items. Generative pretrained transformer 4o (GPT-4o), Claude 3.5 Sonnet, and Gemini 2.5 Flash generated item scores and brief rationales using identical Korean zero-shot rubrics. Concordance between clinician ratings was evaluated using ordinal and binary screening metrics. A paired Wilcoxon test compared patient-level total symptom burden. Binary mismatches were clinically adjudicated as ambiguous or definite overestimation or underestimation with an 8-etiology taxonomy. Model-generated rationales were further meta-evaluated using GPT-5.4, Claude Sonnet 4.6, and Gemini Pro 3.1 across 4 dimensions: citation (use of quoted supporting statements), structure (logical organization of the rationale), mapping (consistency between the rationale and the assigned item rating), and expansion (degree of interpretive elaboration beyond the explicit transcript content). The associations between meta-evaluation results and absolute error were examined using cross-classified mixed-effects models. Results: Out of 101 participants, 88 (87.1%) were predominantly female and had breast cancer as the primary cancer type (n=70, 69.3%). Across 3931 item-level ratings, all models showed good agreement with clinicians (intraclass correlation coefficient: 0.816-0.872), with GPT-4o showing the highest agreement. Claude 3.5 and Gemini 2.5 yielded significantly higher patient-level symptom burden (adjusted Conclusions: In this exploratory clinician-benchmarked evaluation of authentic interviews from a Korean psycho-oncology sample comprising predominantly women and patients with breast cancer, LLMs showed high aggregate concordance with psycho-oncologists' item-level ratings while differing in their error profiles. These findings support further evaluation of LLM-based item-level symptom-rating approaches in psycho-oncology. Validation in larger, more diverse, and independent cohorts is needed.

Indexed as

Large Language ModelsNeoplasmsPsycho-OncologyAdultAgedFemaleHumansMaleMiddle AgedRepublic of Koreadistresslarge language modelmental healthpatients with cancerpsycho-oncology

Identifiers

PMID42715407
PMCPMC13557317

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.