Electronic Thesis/Dissertation
 

Evaluation and prediction of the association consistency between two high-throughput two-sample gene expression datasets

Open Access

Similarity coefficients such as the Jaccard index and Dice similarity coefficient have been widely used in biomedical studies. They are useful in evaluating the level of overlap between two lists of observations (or analysis summaries). To our knowledge, their statistical properties have not been well investigated. Furthermore, we would also like to perform a comparison between two consistency results based on the given formula of similarity coefficient. Based on a general form of similarity coefficient, we have studied the related bias and variance, which can be useful for our selection of a similarity coefficient in practice. Particularly, when the ratio of intersection over the union between two lists of observations (or analysis summaries) is relatively small, the Jaccard index is preferred. Based on our study results for the bias and variance of a similarity coefficient, we have also studied the statistical properties for a comparison test on two similarity evaluations, particularly for the power of the test. For the two-sample test comparing independent consistency results, the similarity coefficient with a smaller value and a larger variance or the similarity coefficient with a larger value and a smaller variance is preferred based on the performance of power. While for the two-sample test comparing dependent consistency results, we can decide the preference of similarity coefficient based on the calculated power according to different datasets. We have discussed the difference between bootstrapping original samples and bootstrapping paired z-scores when conducting the confidence intervals for estimator. We have also demonstrated that a bias-correction procedure is necessary for the confidence interval we conducted when bootstrapping the original samples for a similarity coefficient.Furthermore, when there is only one list of observations (or analysis summaries), it is not feasible to explicitly calculate the similarity coefficient. With only one data set, splitting data randomly can be considered so that two lists of observations (or analysis summaries) can be obtained for the calculation of a similarity coefficient. However, we often prefer to perform such a consistency evaluation without randomly splitting data. Therefore, we have constructed a normal-mixture model with the consideration of kernel density functions to achieve this purpose. We have conducted simulation studies that confirm the asymptotic assumption is consistent and the derivation of formulas is accurate. The simulation results also clearly illustrated the theoretical study results. For applications, we have performed analyses based on the microarray data and RNA sequencing (RNA-seq) data for the Breast invasive carcinoma (BRCA) collected in The Cancer Genome Atlas (TCGA). Based on our application results, RNA-seq is clearly preferred over microarray for its improved similarity evaluation performance. We have also performed analyses based on the RNA-seq data for Colon adenocarcinoma (COAD), Stomach adenocarcinoma (STAD), and Kidney renal clear cell carcinoma (KIRC) studies collected in TCGA. In the two-sample test comparison between these three datasets, Dice is preferred for a larger power.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of guo_gwu_0075A_15870.pdf guo_gwu_0075A_15870.pdf 2022-10-04 Open Access