Advanced Statistical Models in Cancer Research with Omics Data
Open AccessRecent developments of high-throughput technology have made it possible to generate various kinds of omics data, such as mutations, copy number variants, DNA methylation data, and gene expression, etc., which provides us opportunities to better solve the problems in cancer researches, such as cancer subtyping or candidate cancer genes identification. In order to better solve these problems, we need to fully investigate the omics data through statistical models and biological analyses and thus thoroughly understand the effects of omics data. In this dissertation, we developed advanced statistical models to address the problems arising in current cancer researches with omics data. We first proposed an association-signal-annotation boosted similarity network fusion (ab-SNF) method, which adds feature-level association signal annotations as weights aiming to up-weight signal features and down-weight noise features when constructing subject similarity networks using multiple types of omics data jointly and thus boosts the performance in cancer subtyping. In various simulation studies, the proposed ab-SNF outperforms the original similarity network fusion method without weights. Most importantly, the improvement in the subtyping performance due to association-signal-annotation weights is amplified in the integration process. Applications to somatic mutation data, DNA methylation data and gene expression data of three cancer types from The Cancer Genome Atlas (TCGA) project suggest that the proposed ab-SNF method consistently identifies new subtypes in each cancer that more accurately predict patient survival and are more biologically meaningful. Then we proposed mirPLS, a Partial Linear Structure identifier for miRNA data that simultaneously identifies miRNAs of linear or non-linear associations with cancer status when non-linearly associated miRNAs can then be used for subsequent cancer subtyping. Simulation studies showed that mirPLS can identify both non-linearly and linearly outcome-associated miRNAs more accurately than the comparison methods. Using the identified non-linearly associated miRNAs much improves the cancer subtyping accuracy. Applications to miRNA data of three different cancer types from TCGA suggest that the cancer subtypes defined by the non-linearly associated miRNAs identified by mirPLS are consistently more predictive of patient survival and more biological meaningful. We also proposed a Disease-Specific Network Enhancement Prioritization (DiSNEP) framework, which enhances a comprehensive gene network into a disease-specific network using a type of omics data of the disease through a diffusion process. The enhanced disease-specific gene network better reflects true disease gene interactions and improves prioritizing disease-associated genes. In simulations, more genes prioritized by the enhanced disease-specific network are true signals than that by the original gene network or without prioritization. We applied the DiSNEP framework to prioritize cancer-associated gene expression and DNA methylation signal genes for five cancer types using gene-expression-enhanced and DNA-methylation-enhanced cancer-specific networks from TCGA project. We observed that more prioritized candidate genes by the enhanced cancer-specific networks are cancer-related than those by the comparison methods, consistently across all five cancer types considered.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.