Revealing Spam Detection: A Hybrid Machine and Deep Learning Approach to Uncovering Key Textual Features
Open Access DepositedBridging the two problems of achieving a high level of effectiveness in the classification of spam emails and scrutinizing the text of the email that influences the decision, and the classification is the aim of this study. A hybrid machine learning and deep learning model is built to build word matrices of count emails based on punctuation, capitalization, and marked words. Such representations allow the construction of robust automated spam filter systems, wherein system countermeasures can be designed progressively against more advanced spammer intelligence while explaining why certain emails deemed spam were marked as such.The model transforms emails into structured features, while preserving the fundamental structure differentiating spam from non-spam emails. Results with the Spambase dataset indicate that all evaluated models, including Random Forest, Gradient Boosting, MLP, and a Keras-based neural network, achieved high accuracy and F1 scores, confirming spam classification validity. Further analysis also illustrates how recalibrating the estimates from the classification refined the spam probability during the post-processing stage, and feature resilience analyses showed only slight reductions in performance when excising dominant features. All these results demonstrate the enduring effectiveness of this model, which is usable in combative email security systems, and illustrate the challenges posed by emerging spam emails. The framework is able to absorb the impact of gaps in trust and respond to excessive dependence on strict boundaries considering the dynamic nature of email inflow to minimize the adverse effects of spam emails.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.