Electronic Thesis/Dissertation
 

Automation and expansion of the metagenomics analysis methodology using computational tools and statistical methods to support small high-dimensional datasets

Open Access

Advancements of technology have allowed researchers to better understand functional mechanisms of microbial communities in vivo. Here, we provide an overview of how technological advancements has led to current metagenomics analyses practices. Using human and mouse metagenomics data, we were able to modify our metagenomics analysis pipeline to incorporate more automation methods that significantly reduces human error. Later, we dive deeper into downstream analyses method using machine learning. Metagenomics data is high-dimensional and often very difficult and expensive to collect during a clinical trial. Thus, many studies collect publicly available datasets, which contain varying patient characteristics and other uncontrollable factors. As a solution, there should be a general methodology on how to analyze small, high-dimensional metagenomics datasets. In the presence of small sample size, machine learning analysis becomes more difficult to produce reliable predictive models. Throughout this thesis, we discuss how small sample size effects machine learning analysis and discuss current approaches used.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Hopson_gwu_0075M_15556.pdf Hopson_gwu_0075M_15556.pdf 2022-03-06 Open Access