Evidence map›Paper›PMID 41614102›Full record

ArticleFrontiers in psychiatry2025

Leveraging reddit data for context-enhanced synthetic health data generation to identify low self esteem.

Muskan Garg, Xingyi Liu, Eunji Jeon, Joanna M Biernacka, Mark A Frye, Yonas E Geda, Sunghwan Sohn

Abstract read
In one paragraph

Article in Frontiers in psychiatry, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

7 authors.

Muskan GargDepartment of Artificial Intelligence & Informatics, Mayo Clinic, Rochester, MN, United States.
Xingyi LiuDepartment of Artificial Intelligence & Informatics, Mayo Clinic, Rochester, MN, United States.
Eunji JeonDepartment of Artificial Intelligence & Informatics, Mayo Clinic, Rochester, MN, United States.
Joanna M BiernackaDepartment of Quantitative Health Sciences, Mayo Clinic, Rochester, MN, United States.
Mark A FryeDepartment of Psychiatry and Psychology, Mayo Clinic, Rochester, MN, United States.
Yonas E GedaBarrow Neurological Institute, Phoenix, AZ, United States.
Sunghwan SohnDepartment of Artificial Intelligence & Informatics, Mayo Clinic, Rochester, MN, United States.

Funding

Early Detection of Mild Cognitive Impairment, Alzheimer’s Disease and Other Dementias using EHRR01AG068007 · NIA · MAYO CLINIC ROCHESTER · PI Yonas E Geda, Sunghwan Sohn · 2020 to 2026
$3.8M
Advancing women’s care in Alzheimer’s disease and other dementias through EHRRF1AG090341 · NIA · MAYO CLINIC ROCHESTER · PI SOHN, SUNGHWAN · 2025 to 2025
$3.4M
NIA NIH HHS R01 AG068007NIA NIH HHS RF1 AG090341
6 · The paper itself

Abstract

Low self-esteem (LoST) is a latent yet critical psychosocial risk factor that predisposes individuals to depressive disorders. Although structured tools exist to assess self-esteem, their limited clinical adoption suggests that relevant indicators of LoST remain buried within unstructured clinical narratives. The scarcity of annotated clinical notes impedes the development of natural language processing (NLP) models for its detection. Manual chart reviews are labor-intensive and large language model (LLM)-driven (weak) labeling raises privacy concerns. Past studies demonstrate that NLP models trained on LLM-generated synthetic clinical notes achieve performance comparable to, and sometimes better than those trained on real notes. This highlights synthetic data's utility for augmenting scarce clinical corpora while reducing privacy concerns. Prior efforts have leveraged social media data, such as Reddit, to identify linguistic markers of low self-esteem; however, the linguistic and contextual divergence between social media and clinical text limits the generalizability of these models. To address this gap, we present a novel framework that generates context-enhanced synthetic clinical notes from social media narratives and evaluates the utility of small language models for identifying expressions of low self-esteem. Our approach includes a mixed-method evaluation framework: (i) structure analysis, (ii) readability analysis, (iii) linguistic diversity, and (iv) contextual fidelity of LoST cues in source Reddit posts and synthetic notes. This work offers a scalable, privacy-preserving solution for synthetic data generation for early detection of psychosocial risks such as LoST and demonstrates a pathway for translating mental health signals in clinical notes into clinically actionable insights, thereby identifying patients at risk.

Indexed as

clinical notesllamaself-esteemsmall language modelsynthetic data generation

Identifiers

PMID41614102
PMCPMC12847414

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.