A Novel Approach in Financial Fraud Detection: Addressing Data Imbalance with Prompt Engineering, Leveraging Large Language Model Embeddings for Fraud Scoring, and Cost-Effectiveness Optimization
Open Access DepositedThis research examines novel methodologies that leverage large language models (LLMs) for synthetic data generation and embedding tasks to enhance the effectiveness of financial fraud detection. The study also introduces an optimized approach for balancing false negatives and false positives, with a focus on cost-effectiveness, ultimately supporting users in making practical, data-driven decisions.Fraudulent transaction data is often highly imbalanced, necessitating specialized data processing methods to address this challenge. This study employs 14 traditional machine learning and deep learning classifiers alongside six data processing methods to develop fraud detection models and identify the most effective traditional approach for detecting fraudulent transactions. Building on this foundation, the research integrates advanced techniques such as OpenAI, BERT, LLaMA, Gemini, and Gemma to create hybrid models that incorporate categorical variable embeddings within machine learning classifiers. The objective is to enhance fraud detection capabilities beyond traditional methods, extending the work of Bakumenko et al. (2024) (Bakumenko, Hlaváčková-Schindler , Plant , & Hubig, 2024). Furthermore, the study introduces a novel hybrid oversampling technique that combines LangChain, GPT-4o-Mini, and advanced oversampling methods to address data imbalance and improve model performance. To optimize model accuracy, a customized metric is developed to fine-tune hyperparameters by balancing the trade-off between false positives and false negatives. This research advances fraud detection methodologies through three key innovations: 1) Development of hybrid models leveraging LLM-based text embedding, 2) Creation of a hybrid oversampling technique incorporating LLM-based few-shot prompted synthetic data generation, 3) Implementation of a cost-effective optimization strategy to balance false positives and false negatives effectively. Another aspect of this research focuses on both commercial API and publicly licensed light-weighted large language models. This research aims to find the best-in-class methodology to improve financial fraud detection ability within the cybersecurity domain.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.