Electronic Thesis/Dissertation
 

Some Research Progress in Generative Modeling and the Related Applications to Single-Cell RNA-Sequencing and Spatially Resolved Transcriptomics Data

Open Access

Generative modeling is a major statistical learning approach that has found suc- cesses in numerous application fields including natural language processing, computer vision and bioinformatics. Overall, generative models aim at approximating the data- generating process of the training data, and they can generate new synthetic data by sampling from the learned data distribution. In this study, we discuss theoretical properties and application advances of generative models. On the theoretical side, we explored some theoretical properties of Vanilla GAN, the original framework of generative adversarial networks (GANs). On the application side, we studied some recent imputation methods for single-cell RNA-sequencing (scRNA-seq) and spatially resolved transcriptomics (SRT) data. A generative model can randomly generate new plausible values with variations. For imputation methods based on generative modeling, we proposed to apply the multiple imputation (MI) approach by pooling analysis results obtained from multiple randomly imputed values as an effort to improve follow-up analysis of scRNA-seq and SRT data.GANs is a family of generative models. A GAN model trains a generator (G) and a discriminator (D) in a two-player minimax game. Under unlimited capacity for both G and D, in theory, a trained generator of Vanilla GAN can transform random noise to synthetic samples that follow the real data distribution, if the iterative training process converges as expected. However, in practice, the unlimited capacity assumption is rarely satisfied, and the iterative training process is known to be likely unstable. In our study, we formulated a specific family of limited discriminators. Then, we discussed the optimal generated data distribution and stability of the iterative training process (under this limited discriminator formulation). More specifically, based on the proposed discriminator formulation, we provided theoretical conditions under which the training process would move in the expected direction in each iteration. Our results could provide the related mathematical insights for the use of Vanilla GAN.In an scRNA-seq dataset, an excessive number of zeroes (usually considered as missing values) can be frequently observed. This phenomenon often hinders a follow-up analysis of scRNA-seq data. To address this issue, many imputation approaches have been proposed. To our knowledge, most existing imputation methods only generate a one-time imputation for each missing value. Multiple imputation (MI) is a widely accepted approach to address the uncertainty of one-time imputation. There has been a lack of investigation on benefits of applying MI in scRNA-seq data analysis. If an imputation method can generate multiple random imputations, we proposed to apply the MI approach to improve follow-up analysis after imputation. In an application analysis, we implemented MI procedures for clustering analysis, differential expression analysis and pseudotime inference based on a GAN-based imputation method and demonstrated that MI yielded improved analysis results versus one-time imputation.For SRT data at single-cell resolution, in many situations, a proportion of genes cannot be detected due to technical limitations. A collection of methods has been developed to impute expression values of undetected genes in SRT data by borrowing information from corresponding scRNA-seq data on the same tissue. To our knowledge, most SRT imputation methods only provide a single imputed value for each undetected gene. In an application analysis on 14 pairs of publicly available scRNA-seq and SRT datasets, we utilized a well-recognized deep generative imputation method to generate multiple random imputations for undetected genes, estimated gene-gene correlations for each random imputation, and combined multiple estimates into a final estimation of undetected gene-gene correlations. We demonstrated that this MI approach consistently generated better estimation of gene-gene correlation than an estimation based on one-time imputation.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Zhu_gwu_0075A_16338.pdf Zhu_gwu_0075A_16338.pdf 2023-11-14 Open Access