Electronic Thesis/Dissertation
 

Statistical Analysis of DNA Copy Number Variation with Sequencing Data

Open Access

Chromosomal gains and losses comprise an important type of genetic change in human tumors. There are various types of chromosomal alterations, many of which play a role in the initiation and progression of the disease. Recent advances in sequencing technology provide an alternative to DNA microarrays for studying a variety of genomic features, including the DNA copy number variation (CNV) profile of a cancer genome. Cancer occurs from somatic alterations in key genes, including point mutation, copy-number alterations and structural rearrangements. The sequencing data based CNV detection has been explored during the recent years. Most of the existing computational and statistical methods are based on the count type data collected by a sliding window approach.Here we propose to study and analyze copy number variation using distance instead of count type data. A distance can be measured in base pairs (bps) between two adjacent alignments of short reads mapped to a reference sequence. Therefore, our approach does not require a sliding window for scanning the number of short reads. We found that the empirical distribution of the distance data, when applied to a sequencing data set collected for a cancer study (Chiang et al., 2009), could be fitted by a mixture of geometric distributions. Furthermore, our exploration of the distance data revealed a small proportion of observations with distance greater than 5000 bps. An artificial censor of these potential outliers could improve our model fitting. Therefore, we proposed to fit the distance data using a mixture of right censored geometric distributions. Our model can be estimated by the well-established Expectation-Maximization (EM) algorithm. The number of mixture components in our model can be determined by the parametric bootstrap procedure. Our simulation study results showed a satisfactory estimation performance and also a satisfactory hypothesis testing power when the sample size was relatively large. We extended the mixture of generalized linear models (GLMs) to the mixture of right censored geometric distributions. Our extension was to address the issue of GC-content bias. The model was developed based on the Newton-Raphson algorithm as well as the Expectation-Maximization algorithm. Our simulation results indicated a satisfactory estimation performance for relatively large sample size. We then modeled the sequencing data, collected for a cancer study (Chiang et al. 2009), using a mixture GLMs where the distance was considered as the response variable and the GC-content was considered as the predictor. We applied our method to the sequencing data set where we picked a normal and a cancer cell line for a change-points analysis. The rank based inverse normal transformation was applied to the predicted values from the fitted mixture of GLMs and then a recursive combination algorithm was used to detect chromosomal CNV using the transformed data.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Biswas_gwu_0075A_12058.pdf Biswas_gwu_0075A_12058.pdf 2018-01-16 Open Access