Electronic Thesis/Dissertation
 

Explainable Learning with Meaningful Perturbations

Open Access

Deep Neural Networks (DNNs) have transformed our understanding of how sophisticated decision rules can be learned from data. Trained with millions or billions of parameters, deep learning models can produce highly accurate results, sometimes surpassing human expert performance. However, a substantial gap remains between high-accuracy models and poor interpretability of decision-making mechanisms. Most deep predictive models are black-boxes that are challenging to understand even by machine learning experts, not to mention varied domain scientists who are not necessarily well-trained in machine learning. Without interpretation, it is unclear what are the critical features of DNNs, or whether DNNs would even exploit these critical features as humans do. This dissertation establishes research frameworks that bridge this gap by the study of how to interpret increasingly complex DNN models to boost understanding of deep predictive models. This dissertation presents two frameworks for explaining DNNs: (1) design a model-agnostic saliency estimation (MASE) framework to explain DNNs in Natural Language Processing with meaningful perturbations on the embedding space; (2) design a Learnable Mask (LEMA) attribution method for explaining Computer Vision models by optimizing a tunable mask with Gaussian perturbations. Part I: Despite the success of deep neural networks in various areas, their lack of interpretability remains a challenge, particularly in Natural Language Processing (NLP). Most interpretation models for NLP require a deep understanding of the internal structure of neural networks, which can be difficult or even impossible for many users. To address this issue, we present the Model-agnostic Saliency Estimation (MASE) framework, which provides local explanations for predictive models in black-box settings. Unlike other models, MASE requires no prior knowledge of a model’s internal structure and can efficiently estimate input saliency through Normalized Linear Gaussian Perturbations (NLGP). Our experiments demonstrate that MASE outperforms other model-agnostic explaining methods in terms of Delta Accuracy, making it a compelling solution for explaining text-based models. Part II: Attribution-based explainable learning aims to identify the features of an input that significantly influence a model's output. One way to achieve this is by measuring the effects on a model’s output based on small perturbations applied to its input. In this work, we present a novel mask-based attribution method that is computed through perturbation. Unlike existing mask-based methods, which use a fixed-size mask sliding through the input and assume a fixed-size patch to be critical while ignoring the interactions between different patches, this work presents a Learnable Mask (LEMA) attribution approach where the mask size and weights are trainable, and the interactions between different regions are considered. Moreover, the LEMA framework leverages bilinear up-sampling to resize the small mask to the shape of the input image, which is computationally effective. We implemented the LEMA framework to explain a ResNet model trained on ImageNet data. The experimental results demonstrate that LEMA achieves comparable performance as model-specific methods such as CAM, GradCAM, SmoothGradCAM, and performs better than model-agnostic methods such as LIME and Occlusion.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Yang_gwu_0075A_16460.pdf Yang_gwu_0075A_16460.pdf 2023-11-14 Open Access