Electronic Thesis/Dissertation
 

From Hooks to Clicks: A Data-driven Approach to Understanding Language Trends in Phishing Schemes Across Different Attack Vectors

Open Access

Phishing attacks utilize cybersecurity’s significant vulnerabilities, human beings, by leveraging the skills of human deceptiveness and technology to turn their targets into victims. These attacks have become the number one preferred method for adversaries to execute their attacks by successfully understanding the victim’s behavior and emotions to gain their trust. They pose a high threat to organizations in both the business and personal aspects of it. Even though organizations invest in security systems, and updates to their infrastructure they have yet to find a complete defense mechanism against these attacks where individuals follow secure protocols, security awareness training for organizations, and rules and regulations enacted to defend against phishing attacks.With the above understanding in mind, a novel topic to investigate a different form of analysis was drawn out in the open. This research aims to address the critical challenge of identifying phishing attempts within various forms of digital communication, such as emails, text messages, malicious URLs, and job postings. The goal is to leverage Natural Language Processing (NLP) techniques, specifically Bidirectional Encoder Representations from Transformers (BERT), to create a robust model capable of flagging potential phishing content from a diverse range of datasets. The base model introduced is the “PhishNet”. It was further decided to compare the same framework on BART and XLNET models to verify that the BERT approach with the concatenated dataset provides valid credibility to the approach. The model was implemented by assessing a basic predictive model for detecting phishing URLS and building on it by utilizing multi-diverse datasets that were concatenated to address the gaps identified in the predictive models in today’s market. The data was facilitated by different phishing attacks building upon it and combining the data into one form of entry to provide us with a larger scale of understanding of the language used to perform a phishing attack. The phishing dataset shows that there are six main categories of language that are used to grab the attention of the victims – these include obtaining a reward, a form of urgency, curiosity, job requirement, for entertainment purposes and from fear. PhishNet has proven to provide higher accuracy than the current standalone URL blocklist model with an accuracy result of over 99% as it identifies these different psychological traits an attacker might have to perform the attack. With this understanding, we are also able to create a framework for training employees and investigating the feasibility of building the model through a browser extension. The research’s significance lies in offering innovative and effective approaches to prevent and mitigate phishing assaults and strengthening information system security.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Hawamdah_gwu_0075A_16694.pdf Hawamdah_gwu_0075A_16694.pdf 2024-01-11 Open Access