Statistical and Geometric Data Augmentation for Robust Machine Learning-Enabled Decision Support Systems
Open AccessData needed for supervised training of a machine learning and artificial intelligence enabled decision support system is commonly subject to class imbalance and is likely to contain only samples from known categories. Class imbalance can adversely affect the performance of machine learning for prediction and classification. Compounding this problem, a machine learning enabled decision support system may encounter an anomalous pattern from an unknown category that did not originate from the closed set distribution used during training. In this case, decision support system that was trained only on closed set samples may erroneously identify an anomalous pattern as having originated from one of the categories in the closed set, sometimes with very high confidence. This research focuses on two data augmentation techniques to address both class imbalance and the limitations associated with closed set training. First, to address class imbalance, a statistical distance-based approach is formulated which generates samples by both modeling the underlying minority class distribution and by geometrically considering those borderline samples entangled in the majority class. Then, the problem of unknown pattern recognition is considered from a generative perspective in which additional synthetic training samples that represent anomalies are added to the training data. These synthetic samples are generated to optimally balance the desire to place anomalies all along the boundary of the training set in feature space, while not adversely effecting core classification performance on the test set. The efficacy of both approaches is demonstrated across a diverse set of decision support and autonomous applications such as medical diagnostics, character recognition and intrusion detection, and compare its combined classification and identification performance operating on data sets that are subject to class imbalance and which contain anomalous patterns at query time.Finally, data augmentation and synthetic anomaly generation is used in an architectural framework for continuous self-supervised learning. This enables a decision support system to continue to learn and classify new data patterns post-deployment during operation after the pre-deployment process of supervised training.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.