Electronic Thesis/Dissertation
 

Missing Data in Classification and Functional Data Analysis

Open Access

Missing data is a critical issue that poses challenges in both statistics and machine learning research. In Chapter 1, we provided an overview of the missing data problem and the existing solutions. Despite the development of various approaches over time, the current literature still has some limitations. In particular, two essential aspects have received relatively less attention: handling missing functional data and addressing missing data in the context of class imbalance. To address these gaps, this dissertation concentrates on two distinct research projects.Chapter 2 focuses on extending the Inverse Probability Weighting (IPW) and Augmented Inverse Probability Weighting (AIPW) methods to handle missing functional data. This research is motivated by the EMBARC (Establishing Moderators and Biosignatures of Antidepressant Response for Clinical Care for Depression) study, where the goal is to examine the relationship between the pointwise mean current source density (CSD) curve and treatment/response status. By incorporating missing data mechanisms specific to functional data, we aim to provide valid and robust estimates in this context.Chapter 3 introduces the integration of data augmentation into the Multiple Imputation by Chained Equations (MICE) algorithm for imputing missing data. This approach is motivated by the State Impatient Database (SID), where missing data with class imbalance is prevalent. By leveraging data augmentation techniques within the MICE framework, we aim to improve the imputation performance, particularly for imbalanced categorical variables. This project addresses the need for more effective imputation strategies in the presence of class imbalance.By addressing these research gaps, this dissertation contributes to the literature on missing data handling, particularly in the context of missing functional data and class imbalance. The proposed methods and approaches applied to real-world datasets, such as the EMBARC study and the SID database, to demonstrate their effectiveness and practical utility.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Zhang_gwu_0075A_16558.pdf Zhang_gwu_0075A_16558.pdf 2023-11-14 Open Access