Evidence map›Paper›PMID 34169232›Full record

ArticleJAMIA open2021

Automatic gender detection in Twitter profiles for health-related cohort studies.

Yuan-Chi Yang, Mohammed Ali Al-Garadi, Jennifer S Love, Jeanmarie Perrone, Abeed Sarker

Abstract read
In one paragraph

Article in JAMIA open, 2021. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 7 papers.

0numbers the graph read from it
0cells of the map it votes in
7citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

7 citing papers in PubMed.

  1. Article
  2. Article
  3. Can accurate demographic information about people who use prescription medications nonmedically be derived from Twitter?Proceedings of the National Academy of Sciences of the United States of America · 2023
    Article
  4. Article
  5. Article
  6. Article
  7. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Yuan-Chi YangDepartment of Biomedical Informatics, School of Medicine, Emory University, Atlanta, Georgia, USA.
Mohammed Ali Al-GaradiDepartment of Biomedical Informatics, School of Medicine, Emory University, Atlanta, Georgia, USA.
Jennifer S LoveDepartment of Emergency Medicine, School of Medicine, Oregon Health & Science University, Portland, Oregon, USA.
Jeanmarie PerroneDepartment of Emergency Medicine, Perelman School of Medicine, University of Pennsylvania, Philadelphia, Pennsylvania, USA.
Abeed SarkerDepartment of Biomedical Informatics, School of Medicine, Emory University, Atlanta, Georgia, USA.

Funding

Mining Social Media Big Data for Toxicovigilance: Automating the Monitoring of Prescription Medication Abuse via Natural Language Processing and Machine Learning MethodsR01DA046619 · NIDA · UNIVERSITY OF PENNSYLVANIA · PI SARKER, ABEED H · 2018 to 2021
$1.4M
NIDA NIH HHS R01 DA046619
6 · The paper itself

Abstract

objectiveBiomedical research involving social media data is gradually moving from population-level to targeted, cohort-level data analysis. Though crucial for biomedical studies, social media user's demographic information (eg, gender) is often not explicitly known from profiles. Here, we present an automatic gender classification system for social media and we illustrate how gender information can be incorporated into a social media-based health-related study. MATERIALS AND

methodsWe used a large Twitter dataset composed of public, gender-labeled users (Dataset-1) for training and evaluating the gender detection pipeline. We experimented with machine learning algorithms including support vector machines (SVMs) and deep-learning models, and public packages including M3. We considered users' information including profile and tweets for classification. We also developed a meta-classifier ensemble that strategically uses the predicted scores from the classifiers. We then applied the best-performing pipeline to Twitter users who have self-reported nonmedical use of prescription medications (Dataset-2) to assess the system's utility. RESULTS AND DISCUSSION: We collected 67 181 and 176 683 users for Dataset-1 and Dataset-2, respectively. A meta-classifier involving SVM and M3 performed the best (Dataset-1 accuracy: 94.4% [95% confidence interval: 94.0-94.8%]; Dataset-2: 94.4% [95% confidence interval: 92.0-96.6%]). Including automatically classified information in the analyses of Dataset-2 revealed gender-specific trends-proportions of females closely resemble data from the National Survey of Drug Use and Health 2018 (tranquilizers: 0.50 vs 0.50; stimulants: 0.50 vs 0.45), and the overdose Emergency Room Visit due to Opioids by Nationwide Emergency Department Sample (pain relievers: 0.38 vs 0.37).

conclusionOur publicly available, automated gender detection pipeline may aid cohort-specific social media data analyses (https://bitbucket.org/sarkerlab/gender-detection-for-public).

Indexed as

gender detectionmachine learningnatural language processingtoxicovigilanceTwitteruser profiling

Identifiers

PMID34169232
PMCPMC8220305

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.