Electronic Thesis/Dissertation
 

Statistically Consistent Interpretation of Deep Learning

Open Access

Deep neural networks (DNNs) have achieved exceptional performance in many fields including Computer Vision (CV) and Natural Language Processing (NLP). However, the complex structure and black-box setting make these models hard to interpret, and hence their usage is limited. Some methods are proposed to generate interpretations for DNNs based on game theory or local gradient which lacks statistical guarantees. We propose a new framework to estimate the saliency of the input as the local interpretation for the DNNs in both fields, which is computationally efficient and statistically consistent. We define a new estimand as a proxy for the local gradient of pixels for the prediction (i.e. saliency) in Convolutional Neural Networks (CNNs) based on input perturbation. We introduce an efficient estimate using total variation penalization. We prove that our proposed estimand can be estimated efficiently by solving a linear program. Under some regularity conditions, we present its finite-sample convergence rate and propose a novel perturbation scheme to achieve faster convergence than independent Gaussian perturbation. For NLP models, we take the encoding layer as prior information and use normalized Gaussian perturbation to improve the estimation accuracy of the local interpretation. For both studies, experimental results demonstrate that our proposed methods give more plausible interpretations and are much more accurate at identifying key components of input in the black-box setting. In addition, the sensitivity of proposed methods to the target model is verified through a sanity check.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Luo_gwu_0075A_15910.pdf Luo_gwu_0075A_15910.pdf 2022-10-04 Open Access