Electronic Thesis/Dissertation
 

High Dimensional Classification and Clustering with Feature Selection

Open Access

High-dimensional data are usually accompanied by a great number of noise features that contain limited information, which could undermine the classification process. We propose marginal screening and pairwise screening of the variables to select the relevant features. The marginal screening process uses tests of equality of marginal distribution functions and pairwise screening tests the equality of joint distribution functions to select variables that carry discriminating information. The tests use the dissimilarity indices MADD and MADMD, which take advantage of the distance concentration phenomenon in high-dimensional space. We show classification based on these indices have misclassification rates that tend to zero as the number of features (p) diverges.Popular clustering algorithms such as k-means clustering tends to deteriorate in high dimension, low sample size (HDLSS) situations with misclassification rate of 1/2 as p diverges. We propose several clustering algorithms based on the dissimilarity indices MADD and MADMD, such as k-Means, test-based algorithms, and minimal spanning tree. We present several simulations as well as analysis of a real data set to demonstrate the suitability of these techniques in the HDLSS setup.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Wang_gwu_0075M_16373.pdf Wang_gwu_0075M_16373.pdf 2023-11-14 Open Access