Electronic Thesis/Dissertation
 

Machine Learning Approaches For Genomic Sequence Analysis

Open Access Deposited

From Variant-based Prediction To Genomic Language Models

This dissertation advances machine learning (ML) methodologies for genomic sequence analysis across three complementary scales

variant-level associations, evolutionary and functional patterns in whole genomes, and taxonomic classification using 16S rRNA as a marker gene. As genomic data generation accelerates, the challenge lies not only in data volume but in extracting meaningful biological insights from complex, high-dimensional datasets where genotype-phenotype relationships are often nonlinear and obscured by technical noise.The first study introduces deepBreaks, an open-source framework for identifying sequence positions associated with phenotypic traits through comparative ML. By evaluating multiple algorithms and prioritizing variants based on best-fit models, deepBreaks addresses key challenges in genomic association studies, including feature collinearity, high dimensionality, and sequencing noise, providing researchers with an accessible tool for genotype-phenotype investigations. The second study explores genomic language modeling (gLM) as a paradigm for capturing information encoded in DNA sequences. Through systematic investigation of tokenization strategies, model architectures, and training approaches, we developed seqLens, a family of transformer-based models pretrained on datasets spanning over 180 billion nucleotides from prokaryotic and eukaryotic genomes. Comprehensive benchmarking across 19 phenotypic prediction tasks demonstrates that disentangled attention architectures with relative positional encoding, combined with evolutionarily relevant pretraining data, significantly outperform existing approaches. Our analysis provides critical insights into the effects of vocabulary size, domain adaptation strategies, different pooling strategies, and the capacity of gLMs to learn evolutionary relationships. The third study presents 16SgLM, a suite of specialized gLMs for genus-level classification of 16S rRNA sequences from long-read amplicon data. By implementing bootstrap inference with controlled sequence perturbations, 16SgLM provides probabilistic predictions with uncertainty quantification, addressing a fundamental limitation of traditional alignment-based methods. Validation across simulated data, mock communities, and real human microbiome samples demonstrates superior or comparable performance to established tools while offering database-independent classification and confident predictions for sequences that conventional methods leave unclassified. Collectively, these studies demonstrate how ML approaches, from interpretable variant detection to self-supervised language modeling, can address diverse challenges in genomic sequence analysis, providing computational tools that enhance our ability to decode biological information across evolutionary timescales and taxonomic diversity.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Baghbanzadeh_gwu_0075A_17677.pdf Baghbanzadeh_gwu_0075A_17677.pdf 2026-02-26 Open Access