Electronic Thesis/Dissertation
 

Heterogeneous Block Covariance Model for Community Detection & Standardization of Continuous and Categorical Covariates in Sparse Penalized Regressions

Open Access

Heterogeneous Block Covariance Model for Community DetectionCommunity detection is a clustering method based on objects’ pairwise relationships such that objects classified in the same group are more densely connected than objects from different groups. Most of the model-based community detection methods such as the stochastic block model and its variants are designed for networks with binary (yes/no) edges, while in many practical scenarios the edges would have continuous weights, reflecting different degrees of connectivity. The heterogeneous block covariance model (HBCM) imposes a novel clustering structure on the covariance matrix where edges have signed and continuous weights. Furthermore, it takes into account the heterogeneity of objects when forming connections with other objects within a community. A novel variational expectation-maximization (EM) algorithm is proposed to estimate the group membership. The HBCM gives provable consistent estimation of clustering memberships and its superior performance is observed in numerical simulations with different setups. The model is applied to a yeast gene expression dataset to detect the gene clusters regulated by different transcript factors during the yeast cell cycle.Standardization of Continuous and Categorical Covariates in Sparse Penalized RegressionsIn sparse penalized regressions, candidate covariates of different units need to be standardized beforehand so that the coefficient sizes are directly comparable and reflect their relative impacts, which leads to fairer variable selection. However, when covariates of mixed data types (e.g. continuous, binary or categorical) exist in the same dataset, the commonly used standardization methods may lead to different selection probabilities even when the covariates have the same impact on or level of association with the outcome. In the paper, we propose a novel standardization method that targets at generating comparable selection probabilities in sparse penalized regressions for continuous, binary or categorical covariates with the same impact. We illustrate the advantages of the proposed method in simulation studies, and apply it to the National Ambulatory Medical Care Survey data to select factors related to the opioid prescription in US.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Li_gwu_0075A_16069.pdf Li_gwu_0075A_16069.pdf 2022-10-04 Open Access