Evidence map›Paper›PMID 40325286›Full record

ArticleBJC reports2025

Genomics and integrative clinical data machine learning scoring model to ascertain likely Lynch syndrome patients.

Ramadhani Chambuso, Takudzwa Nyasha Musarurwa, Alessandro Pietro Aldera, Armin Deffur, Hayli Geffen, Douglas Perkins, Raj Ramesar

Abstract read
In one paragraph

Article in BJC reports, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Ramadhani ChambusoDepartment of Global Health and Population, Harvard T. Chan School of Public Health, Boston, MA, USA. rchambuso@hsph.harvard.edu.
Takudzwa Nyasha MusarurwaUCT/MRC Genomics and Precision Medicine Research Unit, Division of Human Genetics, Department of Pathology, University of Cape Town, Cape Town, South Africa.
Alessandro Pietro AlderaUCT/MRC Genomics and Precision Medicine Research Unit, Division of Human Genetics, Department of Pathology, University of Cape Town, Cape Town, South Africa.
Armin DeffurUCT/MRC Genomics and Precision Medicine Research Unit, Division of Human Genetics, Department of Pathology, University of Cape Town, Cape Town, South Africa.
Hayli GeffenDepartment of Public Health and Bioinformatics, University of Cape Town, Cape Town, South Africa.
Douglas PerkinsDepartment of Global Health, School of Medicine, University of New Mexico, Albuquerque, NM, USA.
Raj RamesarUCT/MRC Genomics and Precision Medicine Research Unit, Division of Human Genetics, Department of Pathology, University of Cape Town, Cape Town, South Africa.

Funding

Partnership for Global Health Research Training Program (Renewal)-Supplement for OBSSR, NCI, NCI-AIDS and NIDCDD43TW010543 · FIC · HARVARD UNIVERSITY D/B/A HARVARD SCHOOL OF PUBLIC HEALTH · PI Wafaie W Fawzi, DAVIDSON HOWES HAMER · 2017 to 2026
$12.1M
TRAINING AND RESEARCH ON SEVERE MALARIAL ANEMIAD43TW005884 · FIC · UNIVERSITY OF PITTSBURGH AT PITTSBURGH · PI Collins Ouma, Douglas Jay Perkins · 2002 to 2026
$5.2M
FIC NIH HHS D43 TW005884FIC NIH HHS D43 TW010543HBNU Consortium 471804 UCT 37071
6 · The paper itself

Abstract

backgroundLynch syndrome (LS) screening methods include multistep molecular somatic tumor testing to distinguish likely-LS patients from sporadic cases, which can be costly and complex. Also, direct germline testing for LS for every diagnosed solid cancer patient is a challenge in resource limited settings. We developed a unique machine learning scoring model to ascertain likely-LS cases from a cohort of colorectal cancer (CRC) patients.

methodsWe used CRC patients from the cBioPortal database (TCGA studies) with complete clinicopathologic and somatic genomics data. We determined the rate of pathogenic/likely pathogenic variants in five (5) LS genes (MLH1, MSH2, MSH6, PMS2, EPCAM), and the BRAF mutations using a pre-designed bioinformatic annotation pipeline. Annovar, Intervar, Variant Effect Predictor (VEP), and OncoKB software tools were used to functionally annotate and interpret somatic variants detected. The OncoKB precision oncology knowledge base was used to provide information on the effects of the identified variants. We scored the clinicopathologic and somatic genomics data automatically using a machine learning model to discriminate between likely-LS and sporadic CRC cases. The training and testing datasets comprised of 80% and 20% of the total CRC patients, respectively. Group regularisation methods in combination with 10-fold cross-validation were performed for feature selection on the training data.

resultsOut of 4800 CRC patients frorm the TCGA datasets with clinicopathological and somatic genomics data, we ascertained 524 patients with complete data. The scoring model using both clinicopathological and genetic characteristics for likely-LS showed a sensitivity and specificity of 100%, and both had the maximum accuracy, area under the curve (AUC) and AUC for precision-recall (AUCPR) of 1. In a similar analysis, the training and testing models that only relied on clinical or pathological characteristics had a sensitivity of 0.88 and 0.50, specificity of 0.55 and 0.51, accuracy of 0.58 and 0.51, AUC of 0.74 and 0.61, and AUCPR of 0.21 and 0.19, respectively.

conclusionsSimultaneous scoring of LS clinicopathological and somatic genomics data can improve prediction and ascertainment for likely-LS from all CRC cases. This approach can increase accuracy while reducing the reliance on expensive direct germline testing for all CRC patients, making LS screening more accessible and cost-effective, especially in resource-limited settings.

Identifiers

PMID40325286
PMCPMC12053672

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.