Evaluation and prediction of the association consistency between two high-throughput two-sample gene expression datasets
Open AccessSimilarity coefficients such as the Jaccard index and Dice similarity coefficient have been widely used in biomedical studies. They are useful in evaluating the level of overlap between two lists of observations (or analysis summaries). To our knowledge, their statistical properties have not been well investigated. Furthermore, we would also like to perform a comparison between two consistency results based on the given formula of similarity coefficient. Based on a general form of similarity coefficient, we have studied the related bias and variance, which can be useful for our selection of a similarity coefficient in practice. Particularly, when the ratio of intersection over the union between two lists of observations (or analysis summaries) is relatively small, the Jaccard index is preferred. Based on our study results for the bias and variance of a similarity coefficient, we have also studied the statistical properties for a comparison test on two similarity evaluations, particularly for the power of the test. For the two-sample test comparing independent consistency results, the similarity coefficient with a smaller value and a larger variance or the similarity coefficient with a larger value and a smaller variance is preferred based on the performance of power. While for the two-sample test comparing dependent consistency results, we can decide the preference of similarity coefficient based on the calculated power according to different datasets. We have discussed the difference between bootstrapping original samples and bootstrapping paired z-scores when conducting the confidence intervals for estimator. We have also demonstrated that a bias-correction procedure is necessary for the confidence interval we conducted when bootstrapping the original samples for a similarity coefficient.Furthermore, when there is only one list of observations (or analysis summaries), it is not feasible to explicitly calculate the similarity coefficient. With only one data set, splitting data randomly can be considered so that two lists of observations (or analysis summaries) can be obtained for the calculation of a similarity coefficient. However, we often prefer to perform such a consistency evaluation without randomly splitting data. Therefore, we have constructed a normal-mixture model with the consideration of kernel density functions to achieve this purpose. We have conducted simulation studies that confirm the asymptotic assumption is consistent and the derivation of formulas is accurate. The simulation results also clearly illustrated the theoretical study results. For applications, we have performed analyses based on the microarray data and RNA sequencing (RNA-seq) data for the Breast invasive carcinoma (BRCA) collected in The Cancer Genome Atlas (TCGA). Based on our application results, RNA-seq is clearly preferred over microarray for its improved similarity evaluation performance. We have also performed analyses based on the RNA-seq data for Colon adenocarcinoma (COAD), Stomach adenocarcinoma (STAD), and Kidney renal clear cell carcinoma (KIRC) studies collected in TCGA. In the two-sample test comparison between these three datasets, Dice is preferred for a larger power.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.