Electronic Thesis/Dissertation
 

Censored-Poisson Model Based Approach to The Analysis of RNA-seq Data

Open Access

With the development of sequencing technology, ultra high-throughput sequencing data of RNA (RNA-seq) are being increasingly used in biomedical studies. The RNA-seq observations are usually count-type measurements. In addition to reflecting the expression index of gene or isoform, the observed count data are affected by total volume, gene or exon length, isoform structure and GC content so that they cannot be directly used in a follow-up analysis. Therefore, estimating the expressions of genes or isoforms in a RNA-seq study is a fundamental topic in RNA-seq data analysis. In a two-sample RNA-seq study, differential expression analysis is also a fundamental topic.Recently, methods based on the Poisson and negative binomial distributions have been widely used in practice. However, the variation of RNA-seq count-type observations for one gene (or exon) may span a wide range. Due to the complicated experiment procedures of RNA sequencing, extreme high or low counts may be less reliable. There is a lack of statistical methods to address this critical issue.In this study, we developed a censored-Poisson model approach to the analysis of RNA-seq data. The Poisson distribution will be still used to model the majority of observed counts, however, because the highly or lowly expressed counts are likely to be unreliable, we censor them. According to our experience, we considered counts between 15% percentile (15) and 85% percentile (5000) of counts are more reliable and we censor at 15 and 5000. Three general scenarios are discussed: Scenario 1 concerns analysis of RNA-seq data from one exon/gene; Scenario 2 considers analysis of RNA-seq data from multiple exons (one isoform); and Scenario 3 discusses analysis of RNA-seq data from multiple exons (multiple isoforms). We implemented Newton-Raphson numeric procedure with EM-algorithm to achieve the maximum likelihood estimates (MLEs) for our model parameters (including the estimations of gene expression abundances), which have greatly improved comparing to other methods. For two-sample differential expression analysis (normal vs. tumor samples), we developed a likelihood ratio test approach based on our censored-Poisson models. It is well-known that genomic features like GC content impact the RNA sequencing data. These features can be conveniently included in our models as regression covariates adjustments. Simulation studies were conducted and demonstrate the advantage of our censoring approach. For application illustrations, we applied our methods to the RNA-seq data from The Cancer Genome Atlas (TCGA, breast cancer study). We consider 4 isoforms of gene TP53 for applying the censored-Poisson model and differential expression analysis. The result shows significant difference between normal and tumor subjects with p-value <0.05 based on permutation. We also use the gene SPDYE6 with one transcript for regression analysis. Both normal and tumor samples give similar unimodal pattern between exon counts and GC content based on estimated coefficients.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Chen_gwu_0075A_13349.pdf Chen_gwu_0075A_13349.pdf 2018-01-16 Open Access