Electronic Thesis/Dissertation
 

Detecting Pay Outliers in the United States Labor Force with Context-Aware Attention-Enhanced Latent Diffusion Models

Open Access Deposited

Outlier or anomaly detection is a critical process in many industries. It is an important topic of research in artificial intelligence with applications in fields like fraud detection, cyber-security, and many others. This research focuses on the process of anomaly detection using unsupervised methods (no labels used or available). Primary goal focuses on anomaly detection for compensation but rapidly expands into general anomaly detection. For the initial goal of anomaly detection in compensation, it is recognized as a critical part of any operating businesses as compensation is closely regulated internally and externally. Early accurate detection becomes crucial to identifying potential underlying issues. Traditional methods used like density estimation, inter-quartiles ranges, or simple fixed rules suffer from inefficiencies flagging outliers and generating a larger number of false positives and negatives (He et al., 2005

Pang, Cao, et al., 2021). Most data which need to be reviewed for potential anomalies comes unlabeled or labelling would be too costly or impossible to do within the allowed timeframes, thus there is a need for accurate unsupervised anomaly detection systems. This research introduces CALDM, which stands for Context-Aware Latent Diffusion Model. This method is an unsupervised method for anomaly detection (no labels used for training). This method offers advantages in cost and speed. Also, the introduction of attention mechanisms adds a contextual dimension to the anomaly detection process, which is critical as context plays a fundamental role in determining if a sample is anomalous or not. CALDM was tested for its anomaly detection role in compensation using a dataset from the Bureau of Labor Statistics (BLS). For the wider-scope and more general anomaly detection role, CALDM was tested using a publicly available group of datasets used for benchmarking outlier detection systems called ADBench (Han et al., 2022). These datasets were used to compare CALDM against a comprehensive list of commonly used unsupervised anomaly detection algorithms available publicly. The test results showed that CALDM outperformed all other selected unsupervised anomaly detection algorithms. Its ability to add a contextual dimension to the process gave CALDM a leading edge over other systems tested. These results position CALDM as powerful general-purpose algorithm across industries outside the primary and original focus which was compensation.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Falck_gwu_0075A_17640.pdf Falck_gwu_0075A_17640.pdf 2025-12-12 Open Access