Electronic Thesis/Dissertation
 

Towards Trustworthy Deep Learning Models

Open Access Deposited

Deep learning is a subset of artificial intelligence (AI) that relies on multi-layered neural networks to identify patterns and make predictions from complex, unstructured data. As deep learning continues to advance rapidly, the trustworthiness of these models has become an important area of research. While in this thesis we focus on the trustworthiness of newer models---Graph Neural Networks (GNNs) and text-to-image diffusion models---the technology and analysis are applicable to a wide range of deep learning models.Specifically, GNNs are a set of deep learning models working on graph-structured data, whereas text-to-image diffusion models are a class of generative models that produce images conditioned on textual descriptions. This thesis aims to investigate and improve the trustworthiness of these two types of models. In Chapter 1, we introduce the background of trustworthiness of deep learning models, specifically the explainability and privacy in GNNs and prompt inversion in text-to-image diffusion models. We also summarize related works and highlights the challenges and contributions of this dissertation. In Chapter 2, we propose Illuminati, an explanation method that provides a comprehensive explanation for GNNs. By learning the importance scores for both graph structure and node attributes, Illuminati is able to accurately explain the prediction contribution from nodes, edges, and attributes. We apply Illuminati to two cybersecurity applications. Our experiments show Illuminati achieves high explanation fidelity. We also demonstrate the practical usage of Illuminati in cybersecurity applications. In Chapter 3, we design Revelio, a novel method to provide faithful explanations of message flows in GNNs. To our knowledge, it is the most scalable flow-based explanation method currently available. We have shown that Revelio achieves as much as 28x speedup over the state-of-the-art flow-based GNN explanation methods and can scale to graphs larger than those prior works can realistically process. Additionally, this speedup does not come with a tradeoff in faithfulness. In Chapter 4, we present Maui, the first link inference attack that operates under the most challenging security setting. By leveraging the vertical partition of graph structures and node features, Maui is capable of accurately inferring the existence of edges between nodes using only black-box access to an inference API. Our evaluation on six real-world datasets has demonstrated the effectiveness of Maui in varying scopes, even under defense mechanisms. The results of this study highlight the pressing need for new privacy-preserving techniques to defend against GNN privacy attacks. In Chapter 5, we explore hard prompt inversion for text-to-image diffusion models with a focus on interpretability. We propose Inverto for generating feasible, word-level prompts that are both human-readable and semantically meaningful, with potential for extension to multi-model generalization. Our approach produces interpretable prompts that can guide image generation effectively. We also highlight important safety concerns associated with prompt inversion, showing how inversion techniques can inadvertently be used to bypass content filters by exploiting semantic and visual associations. Chapter 6 concludes this thesis and outlines my plans for further enhancing trustworthiness beyond current works.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of He_gwu_0075A_17489.pdf He_gwu_0075A_17489.pdf 2025-12-11 Open Access