Enhancing Insider Threat Detection Using User-Based Sequencing and Transformer Encoders
Open Access DepositedThe increasing prevalence of insider threats and their costly impact have garnered considerable attention from researchers and security practitioners alike. According to the "The Cost of Insider Risk" report published by the Ponemon Institute in 2023, 71% of the surveyed companies reported experiencing between 21 and more than 40 insider incidents per year, representing a 4% increase from the same report in 2022. Unlike external intruders, insiders are more capable of executing attacks, often without advanced technical skills or experience. Detecting these attacks is particularly challenging due to the insider’s authorized access to the system. Many machine-learning detection algorithms researchers have utilized todate have demonstrated varying results in identifying malicious insider activities. Although innovative, those algorithms typically view threat data as snapshots, often missing the long-term user behavior patterns. In this study, we developed a novel User-Based Sequencing (UBS) method to restructure the CERT dataset into a sequential format and use a Transformer Encoder model to learn from normal user behavior. The output from the Transformer is then fed into three ML algorithms, One-Class Support Vector Machine (OCSVM), Local Outlier Factor (LOF), and Isolation Forest (IFOREST), to detect anomalies. Our Transformer model achieved state-of-the-art performance, with a high accuracy rate of 96.61%, 99.43% recall, an F1-score of 96.38%, an AUROC value of 95.00%, and a low False Negative Rate (FNR) and False Positive Rate (FPR) of 0.0057 and 0.0571, respectively. The success of sequence-based models such as transformers is primarily credited to their innate capability to learn from the sequential representations embedded within languages naturally. Our proposed User-Based Sequencing aims to build those same sequential patterns on tabular data to ease the learning of temporal dependencies, thus allowing the modelsto adapt to new latent representations of users’ behavior as their work patterns evolve. User-based sequencing should improve the detection accuracy of sequential models like transformers as well as non-sequential models like autoencoders. This assumption is validated in our experiment when we run the autoencoder on UBS versus tabular data. We saw an increase of 28.28%, 14.38%, and 30.77% in accuracy, precision, and recall, respectively. Thus, this research not only bridges the gap by employing the Transformer Encoders to effectively tackle the insider threat problem but also holds significant implications for future studies in the cybersecurity domain. Our novel User-Based Sequencing structure can serve as a robust foundation for developing advanced machine-learning models to address various cybersecurity threats.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.
| Thumbnail | Title | Date Uploaded | Visibility | Actions |
|---|---|---|---|---|
|
|
elbasheer_gwu_0075A_16998.pdf | 2025-04-09 | Open Access |
|