skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


Title: Cancer Biomarker Discovery from Gene Co-expression Networks Using Community Detection Methods
Finding the network biomarkers of cancers and the analysis of cancer driving genes that are involved in these biomarkers are essential for understanding the dynamics of cancer. Clusters of genes in co-expression networks are commonly known as functional units. This work is based on the hypothesis that the dense clusters or communities in the gene co-expression networks of cancer patients may represent functional units regarding cancer initiation and progression. In this study, RNA-seq gene expression data of three cancers - Breast Invasive Carcinoma (BRCA), Colorectal Adenocarcinoma (COAD) and Glioblastoma Multiforme (GBM) - from The Cancer Genome Atlas (TCGA) are used to construct gene co-expression networks using Pearson Correlation. Six well-known community detection algorithms are applied on these networks to identify communities with five or more genes. A permutation test is performed to further mine the communities that are conserved in other cancers, thus calling them conserved communities. Then survival analysis is performed on clinical data of three cancers using the conserved community genes as prognostic co-variates. The communities that could distinguish the cancer patients between high- and low-risk groups are considered as cancer biomarkers. In the present study, 16 such network biomarkers are discovered.  more » « less
Award ID(s):
1901628 1651917
PAR ID:
10141532
Author(s) / Creator(s):
;
Date Published:
Journal Name:
2019 IEEE International Conference on Bioinformatics and Biomedicine (IEEE BIBM)
Page Range / eLocation ID:
2097 to 2104
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Two graph theoretic concepts—clique and bipartite graphs—are explored to identify the network biomarkers for cancer at the gene network level. The rationale is that a group of genes work together by forming a cluster or a clique-like structures to initiate a cancer. After initiation, the disease signal goes to the next group of genes related to the second stage of a cancer, which can be represented as a bipartite graph. In other words, bipartite graphs represent the cross-talk among the genes between two disease stages. To prove this hypothesis, gene expression values for three cancers— breast invasive carcinoma (BRCA), colorectal adenocarcinoma (COAD) and glioblastoma multiforme (GBM)—are used for analysis. First, a co-expression gene network is generated with highly correlated gene pairs with a Pearson correlation coefficient ≥ 0.9. Second, clique structures of all sizes are isolated from the co-expression network. Then combining these cliques, three different biomarker modules are developed—maximal clique-like modules, 2-clique-1-bipartite modules, and 3-clique-2-bipartite modules. The list of biomarker genes discovered from these network modules are validated as the essential genes for causing a cancer in terms of network properties and survival analysis. This list of biomarker genes will help biologists to design wet lab experiments for further elucidating the complex mechanism of cancer. 
    more » « less
  2. null (Ed.)
    Abstract The ability to predict the efficacy of cancer treatments is a longstanding goal of precision medicine that requires improved understanding of molecular interactions with drugs and the discovery of biomarkers of drug response. Identifying genes whose expression influences drug sensitivity can help address both of these needs, elucidating the molecular pathways involved in drug efficacy and providing potential ways to predict new patients’ response to available therapies. In this study, we integrated cancer type, drug treatment, and survival data with RNA-seq gene expression data from The Cancer Genome Atlas to identify genes and gene sets whose expression levels in patient tumor biopsies are associated with drug-specific patient survival using a log-rank test comparing survival of patients with low vs. high expression for each gene. This analysis was successful in identifying thousands of such gene–drug relationships across 20 drugs in 14 cancers, several of which have been previously implicated in the respective drug’s efficacy. We then clustered significant genes based on their expression patterns across patients and defined gene sets that are more robust predictors of patient outcome, many of which were significantly enriched for target genes of one or more transcription factors, indicating several upstream regulatory mechanisms that may be involved in drug efficacy. We identified a large number of genes and gene sets that were potentially useful as transcript-level biomarkers for predicting drug-specific patient survival outcome. Our gene sets were robust predictors of drug-specific survival and our results included both novel and previously reported findings, suggesting that the drug-specific survival marker genes reported herein warrant further investigation for insights into drug mechanisms and for validation as biomarkers to aid cancer therapy decisions. 
    more » « less
  3. Background: Though the development of targeted cancer drugs continues to accelerate, doctors still lack reliable methods for predicting patient response to standard-of-care therapies for most cancers. DNA methylation has been implicated in tumor drug response and is a promising source of predictive biomarkers of drug efficacy, yet the relationship between drug efficacy and DNA methylation remains largely unexplored. Method: In this analysis, we performed log-rank survival analyses on patients grouped by cancer and drug exposure to find CpG sites where binary methylation status is associated with differential survival in patients treated with a specific drug but not in patients with the same cancer who were not exposed to that drug. We also clustered these drug-specific CpG sites based on co-methylation among patients to identify broader methylation patterns that may be related to drug efficacy, which we investigated for transcription factor binding site enrichment using gene set enrichment analysis. Results: We identified CpG sites that were drug-specific predictors of survival in 38 cancer-drug patient groups across 15 cancers and 20 drugs. These included 11 CpG sites with similar drug-specific survival effects in multiple cancers. We also identified 76 clusters of CpG sites with stronger associations with patient drug response, many of which contained CpG sites in gene promoters containing transcription factor binding sites. Conclusion: These findings are promising biomarkers of drug response for a variety of drugs and contribute to our understanding of drug-methylation interactions in cancer. Investigation and validation of these results could lead to the development of targeted co-therapies aimed at manipulating methylation in order to improve efficacy of commonly used therapies and could improve patient survival and quality of life by furthering the effort toward drug response prediction. 
    more » « less
  4. null (Ed.)
    Ribonuclease (RNase) H2 is a key enzyme for the removal of RNA found in DNA-RNA hybrids, playing a fundamental role in biological processes such as DNA replication, telomere maintenance, and DNA damage repair. RNase H2 is a trimer composed of three subunits, RNASEH2A being the catalytic subunit. RNASEH2A expression levels have been shown to be upregulated in transformed and cancer cells. In this study, we used a bioinformatics approach to identify RNASEH2A co-expressed genes in different human tissues to underscore biological processes associated with RNASEH2A expression. Our analysis shows functional networks for RNASEH2A involvement such as DNA replication and DNA damage response and a novel putative functional network of cell cycle regulation. Further bioinformatics investigation showed increased gene expression in different types of actively cycling cells and tissues, particularly in several cancers, supporting a biological role for RNASEH2A but not for the other two subunits of RNase H2 in cell proliferation. Mass spectrometry analysis of RNASEH2A-bound proteins identified players functioning in cell cycle regulation. Additional bioinformatic analysis showed that RNASEH2A correlates with cancer progression and cell cycle related genes in Cancer Cell Line Encyclopedia (CCLE) and The Cancer Genome Atlas (TCGA) Pan Cancer datasets and supported our mass spectrometry findings. 
    more » « less
  5. Wei, Yanjie; Li, Min; Skums, Pavel; Cai, Zhipeng (Ed.)
    Novel discoveries of biomarkers predictive of drug-specific responses not only play a pivotal role in revealing the drug mechanisms in cancers, but are also critical to personalized medicine. In this study, we identified drug-specific biomarkers by integrating protein expression data, drug treatment data and survival outcome of 7076 patients from The Cancer Genome Atlas (TCGA). We first defined cancer-drug groups, where each cancer-drug group contains patients with the same cancer and treated with the same drug. For each protein, we stratified the patients in each cancer-drug group by high or low expression of the protein, and applied log-rank test to examine whether the stratified patients show significant survival difference. We examined 336 proteins in 98 cancer-drug groups and identified 65 protein-cancer-drug combinations involving 55 unique proteins, where the protein expression levels are predictive of drug-specific survival outcomes. Some of the identified proteins were supported by published literature. Using the gene expression data from TCGA, we found the mRNA expression of ∼11% of the drug-specific proteins also showed significant correlation with drug-specific survival, and most of these drug-specific proteins and their corresponding genes are strongly correlated. 
    more » « less