Evidence map›Paper›PMID 40868395›Full record

ArticleBioengineering (Basel, Switzerland)2025

Supervised Learning and Large Language Model Benchmarks on Mental Health Datasets: Cognitive Distortions and Suicidal Risks in Chinese Social Media.

Hongzhi Qi, Guanghui Fu, Jianqiang Li, Changwei Song, Wei Zhai, Dan Luo, Shuo Liu, Yijing Yu, Bingxiang Yang, Qing Zhao

Abstract read
In one paragraph

Article in Bioengineering (Basel, Switzerland), 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 8 papers.

0numbers the graph read from it
0cells of the map it votes in
8citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

8 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Review
  8. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Hongzhi QiCollege of Computer Science, Beijing University of Technology, Beijing 100124, China.ORCID 0009-0004-8007-3257
Guanghui FuInstitut du Cerveau-Paris Brain Institute-ICM, Sorbonne Université, CNRS, Inria, Inserm, AP-HP, Hôpital de la Pitié-Salpêtrière, 75013 Paris, France.
Jianqiang LiCollege of Computer Science, Beijing University of Technology, Beijing 100124, China.ORCID 0000-0003-1995-9249
Changwei SongCollege of Computer Science, Beijing University of Technology, Beijing 100124, China.
Wei ZhaiCollege of Computer Science, Beijing University of Technology, Beijing 100124, China.
Dan LuoSchool of Nursing, Wuhan University, Wuhan 430071, China.
Shuo LiuSchool of Nursing, Wuhan University, Wuhan 430071, China.
Yijing YuSchool of Nursing, Wuhan University, Wuhan 430071, China.
Bingxiang YangSchool of Nursing, Wuhan University, Wuhan 430071, China.
Qing ZhaoCollege of Computer Science, Beijing University of Technology, Beijing 100124, China.

Funding

Beijing Natural Science Foundation 7254302Fundamental Research Funds for the Central Universities 2042022kf1218 and 2042022kf1037National Natural Science Foundation of China 72174152 and 72474166
6 · The paper itself

Abstract

On social media, users often express their personal feelings, which may exhibit cognitive distortions or even suicidal tendencies on certain specific topics. Early recognition of these signs is critical for effective psychological intervention. In this paper, we introduce two novel datasets from Chinese social media: SOS-HL-1K for suicidal risk classification, which contains 1249 posts, and SocialCD-3K, a multi-label classification dataset for cognitive distortion detection that contains 3407 posts. We conduct a comprehensive evaluation using two supervised learning methods and eight large language models (LLMs) on the proposed datasets. From the prompt engineering perspective, we experiment with two types of prompt strategies, including four zero-shot and five few-shot strategies. We also evaluate the performance of the LLMs after fine-tuning on the proposed tasks. Experimental results show a significant performance gap between prompted LLMs and supervised learning. Our best supervised model achieves strong results, with an F1-score of 82.76% for the high-risk class in the suicide task and a micro-averaged F1-score of 76.10% for the cognitive distortion task. Without fine-tuning, the best-performing LLM lags by 6.95 percentage points in the suicide task and a more pronounced 31.53 points in the cognitive distortion task. Fine-tuning substantially narrows this performance gap to 4.31% and 3.14% for the respective tasks. While this research highlights the potential of LLMs in psychological contexts, it also shows that supervised learning remains necessary for more challenging tasks.

Indexed as

cognitive distortionsdeep learninglarge language modelmental healthsocial mediasuicide detection

Identifiers

PMID40868395
PMCPMC12383806

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.