High Dimensional Classification and Clustering with Feature Selection
Open AccessHigh-dimensional data are usually accompanied by a great number of noise features that contain limited information, which could undermine the classification process. We propose marginal screening and pairwise screening of the variables to select the relevant features. The marginal screening process uses tests of equality of marginal distribution functions and pairwise screening tests the equality of joint distribution functions to select variables that carry discriminating information. The tests use the dissimilarity indices MADD and MADMD, which take advantage of the distance concentration phenomenon in high-dimensional space. We show classification based on these indices have misclassification rates that tend to zero as the number of features (p) diverges.Popular clustering algorithms such as k-means clustering tends to deteriorate in high dimension, low sample size (HDLSS) situations with misclassification rate of 1/2 as p diverges. We propose several clustering algorithms based on the dissimilarity indices MADD and MADMD, such as k-Means, test-based algorithms, and minimal spanning tree. We present several simulations as well as analysis of a real data set to demonstrate the suitability of these techniques in the HDLSS setup.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.