Evidence map›Paper›PMID 40794953›Full record

ArticleBriefings in bioinformatics2025

Towards machine learning fairness in classifying multicategory causes of deaths in colorectal or lung cancer patients.

Catherine H Feng, Fei Deng, Mary L Disis, Nan Gao, Lanjing Zhang

Abstract read
In one paragraph

Article in Briefings in bioinformatics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

5 authors.

Catherine H FengDepartment of Molecular and Cellular Biology, Harvard University, 52 Oxford St, Cambridge, MA, 02138 United States.ORCID 0000-0002-7413-9748
Fei DengDepartment of Chemical Biology, Ernest Mario School of Pharmacy, Rutgers University, 160 Frelinghuysen Rd., Piscataway, NJ 08854, United States.
Mary L DisisUW Medicine Cancer Vaccine Institute University of Washington, 850 Republican St, Seattle, WA 98109, United States.
Nan GaoDepartment of Pharmacology, Physiology, and Neuroscience, New Jersey Medical School, Rutgers University, 185 South Orange Avenue, Newark, NJ 07101, United States.
Lanjing ZhangDepartment of Chemical Biology, Ernest Mario School of Pharmacy, Rutgers University, 160 Frelinghuysen Rd., Piscataway, NJ 08854, United States.ORCID 0000-0001-5436-887X

Funding

Paneth cell heterogeneity in infection and inflammationR01DK132885 · NIDDK · RUTGERS THE STATE UNIV OF NJ NEWARK · PI Nan Gao · 2022 to 2026
$2.4M
Screening and confirmatory machine learning for explainable modeling of non-cancer deaths in cancer patientsR37CA277812 · NCI · RUTGERS BIOMEDICAL AND HEALTH SCIENCES · PI Lanjing Zhang · 2022 to 2026
$1.6M
NCI NIH HHS R37 CA277812NIDDK NIH HHS R01 DK132885NIH HHS R01DK132885NIH HHS R37CA277812U.S. National Science Foundation IIS-2128307
6 · The paper itself

Abstract

Classification of patient multicategory survival outcomes is important for personalized cancer treatments. Machine learning (ML) algorithms have increasingly been used to inform healthcare decisions, but these models are vulnerable to biases in data collection and algorithm creation. ML models have previously been shown to exhibit racial bias, but their fairness towards patients from different age and sex groups have yet to be studied. Therefore, we compared the multimetric performances of five ML models (random forests, multinomial logistic regression, linear support vector classifier, linear discriminant analysis, and multilayer perceptron) when classifying colorectal cancer patients (n = 589) of various age, sex, and racial groups using The Cancer Genome Atlas data. All five models exhibited biases for these sociodemographic groups. We then repeated the same process on lung adenocarcinoma (n = 515) to validate our findings. Surprisingly, most models tended to perform more poorly overall for the largest sociodemographic groups. Methods to optimize model performance, including testing the model on merged age, sex, or racial groups, and creating a model trained on and used for an individual or merged sociodemographic group, show potential to reduce disparities in model performance for different groups. This is supported by our regression analysis showing associations between model choice and methodology used with reduced performance disparities across demographic subgroups. Notably, these methods may be used to improve ML fairness while avoiding penalizing the model for exhibiting bias and thus sacrificing overall performance.

Indexed as

Colorectal NeoplasmsLung NeoplasmsMachine LearningAgedAlgorithmsFemaleHumansMaleMiddle Agedcolorectal cancerlung cancermachine learningmachine learning fairnessmultilabel classificationsurvival

Identifiers

PMID40794953
PMCPMC12342732

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.