Ensemble Machine Learning to Detect Mobile Payment Fraud With User Risk Classification
Open Access DepositedDespite the effectiveness of mobile payments, mobile payment fraud has increased year on year due to a lack of fraud detection tools that keep pace with evolving fraud patterns. This has been seen as the dark side of mobile payment, and has also resulted in financial losses to mobile payment users. This study examines the effectiveness of combining user risk classification and fraud detection methods to combat mobile payment fraud. The study explores different classification methods to classify users based on risk, and it also explores deep learning methods for detecting fraud and supporting users in accurately reducing it.Fraud detection operates with an imbalanced dataset, which reflects the reality on the ground where the fraud samples represent the minority of the rest. In this study, four different data sampling techniques were used along with three distinct methodologies for user risk classification and three methodologies for payment fraud detection to identify the most effective in user risk classification and one that excels in payment fraud detection effectively. Among the user risk classification methodologies, this study is built on the foundation of machine learning models such as Random Forest, XGBoost, and AdaBoost to explore their performance and choose the most effective model. In addition, for the deep learning models, the study is built on the foundation models such as Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), and autoencoders (AEs) models to detect fraud. Furthermore, the study introduced an ensemble machine learning model that incorporates a user risk classifier with the deep learning model to detect payment fraud. Therefore, this research aims to determine which of the four models — CNN, LSTM, AEs, and the novel ensemble model that combines user risk classification and the deep learning method — is most effective in accurately detecting payment fraud. The study showed that among user risk classification methods, Random Forest outperformed the others, achieving 0.9981 accuracy, 0.9996 AUC, and 0.9285 F1 score on the original dataset. In addition, in fraud detection using deep learning methods, CNN outperformed other methods, achieving 0.9660 accuracy and 0.8495 F1 score on the SMOTE-oversampled dataset. In conclusion, the stacked ensemble machine learning model, consisting of a Random Forest user risk classifier and a CNN fraud detection model, outperformed the deep learning models, achieving an AUC of 0.8913, a recall of 0.7919, and an F1 score of 0.8531.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.
| Thumbnail | Title | Date Uploaded | Visibility | Actions |
|---|---|---|---|---|
|
|
Dusengimana_gwu_0075A_17725.pdf | 2025-12-15 | Open Access |
|