Electronic Thesis/Dissertation
 

Development and Evaluation of Cloud Computing Infrastructures for Next-Generation Sequence Data Analysis

Open Access

AbstractBackgroundAfter the completion of the Human Genome Project, there has been a high demand for low cost sequencing. This has given rise to many high throughput Next Generation Sequencing techniques. Following this there has been a steep increase in the amount of sequencing data. This has created some problems for the biologists which include storage of the data, processing of the data and sharing the data. While large genomic institutes like Broad Institute and J Craig Venter Institute have the necessary infrastructure and resource to process these data, the same becomes very difficult for a small laboratory with individual researcher. With the advent of bench top genome sequencers like Miseq from Illumina and GS Junior from Roche, small labs can generate huge amount of data from complete genome sequencing of viral, bacterial and fungal genome in very less time. With this data small labs need additional fund for building clusters, for hiring experts to manage the clusters. With the increase in data they will have to upgrade the infrastructure. It also leads to minimal utilization of the hardware and duplication of data across labs. Cloud computing can be a viable solution to this problem. Researchers can rent computation capacity on demand from companies like Amazon and Google AppEngine which rent computational resource in a pay as you go model. This way they don't have to invest on buying and maintaining any hardware, and they don't have to pay when not using it. In order to make the cloud infrastructure more compatible with biological workflows certain approaches has been made. These approaches include development of Cloud BioLinux which offers an on-demand, cloud computing solution for the bioinformatics community, and is available for use on private or publicly accessible, commercially hosted cloud computing infrastructure such as Amazon EC2 and CloVR which is a desktop application for push-button automated sequence analysis that can utilize cloud computing resources to provide improved access to bioinformatics workflows and distributed computing resource. In this study we use the metagenomic pipeline from the J Craig Venter private cluster to study its behavior on the Amazon EC2 cloud using Cloud BioLinux virtual machine. ResultThe metagenomic pipeline was executed on the Amazon EC2 using the Cloud Biolinux virtual machine with different configurations. First a preliminary experiment was conducted using only one input file. Based on the results of this test the final experiment was designed by making changes to the preliminary experiment. The final experiment was conducted with four input files. Attributes were selected from the usage over the entire cluster and the values were stored in Excel worksheets. From these worksheets, graphs were produces showing the relation of the attributes with the addition of nodes to the cluster. Based on these graphs the behavior of the pipeline was studied and discussed.ConclusionBased on the result of the execution of the pipeline on Amazon EC2 with different configuration the behavior of the pipeline was studied. All the attributes that were tested, only the relevant attributes which affected the efficiency and behavior of the pipeline were discussed. Among these attributes we have some attributes whose values increases as the number of nodes increased, whereas the expected result was a decrease in the value with the increase in nodes. For these attributes it was concluded that, when a small size of data is processed over a smaller size cluster, the overheads and the network latency contribute to the increase in the value. For some attributes the value increased when a bigger master is used as compared to the smaller master. Some attributes are constant across the nodes. Any spike in these attributes can be suggested to be an abnormality in the execution of the pipeline. There were also some attributes which were dependent on the sequences specific to the input file. Using the values of these attributes we can understand the behavior of the pipeline and make changes to it in order to make it more efficient inside the cloud infrastructure.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Sarangi_gwu_0075M_11573.pdf Sarangi_gwu_0075M_11573.pdf 2018-01-16 Open Access