A Data-Driven Approach to Malware Detection in Android Phones
Open Access DepositedEnhancing Network Resilience in the Telecommunication Sector
Abstract of PraxisEnhancing Network Resilience in the Telecommunication Sector
0.963) while maintaining an acceptable model size (99.71 KB) and similar low memory usage, making it the second-best choice for mobile deployments. While traditional models such as Random Forest and XGBoost achieved slightly higher raw accuracy and F1-scores, their larger model sizes limited their suitability for resource-constrained Android environments.
96.31%, F1-score
A Data-Driven Approach to Malware Detection in Android Phones “The telecom sector, a backbone of modern technology, has seen a significant increase in sophisticated malware attacks (EdgeLabs, 2023). The widespread integration of Android devices into telecom networks has heightened the need for robust malware detection mechanisms (Atanassov, 2021). In 2023, Kaspersky reported a 50% increase in mobile malware incidents, totaling 33.8 million attacks, underscoring the urgency for more enhanced detection solutions (Kaspersky, 2024). To tackle the new threat, this study used a two-fold strategy that experimented with both supervised and unsupervised machine learning (ML) models to enhance malware detection and classification inside telecom networks. This study used the Canadian Institute for Cybersecurity dataset, CCS-CIC-AndMal-2020, which entailed Android malware characterized by a total of 145 features (Lashkari, 2020). A wrapper-based feature selection technique was applied to create an optimized set of 10 high-relevancy features to fill redundancy gaps in previous research (Xiao, 2020). The supervised part addressed intelligent learning architecture more specifically, such as the hybrid network of the Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU). These models are efficient in handling sequential patterns in Android traffic while eliminating the problem of vanishing gradients that are widely associated with traditional RNNs (Ayyalasomayajula, 2024). On the other hand, unsupervised models, including Isolation Forest, Autoencoder, and K-Means clustering, were tested to detect unknown threats without using any labeled data (Ayyalasomayajula, 2024). Additionally, a lightweight hybrid miniLSTM+miniGRU model was introduced, which can be utilized in resource-limited contexts, including mobile and edge devices. While the model is effective in malware detection, it exhibits low memory utilization and minimal computational burden, making it ideal for real-time applications. To benchmark the results, twelve (12) models, in addition to their lightweight versions, were used. These models include Random Forest (RF), XGBoost, Support Vector Machine (SVM), Convolutional Neural Network (CNN), LSTM, GRU, RF+MiniLSTM, Isolation Forest, Autoencoder, and K-Means clustering. Both predictive accuracy (e.g., F1-score, AUC) and operational efficiency (model size, inference time, and memory requirements) were used as the evaluation criteria. The results revealed that the Hybrid miniLSTM+miniGRU model provided the best overall balance between predictive accuracy and operational efficiency, achieving a strong accuracy of 95.05%, an F1-score of 0.9506, a compact model size (66.66 KB), and low memory usage (11.6 MB). The Hybrid LSTM+GRU model also demonstrated high predictive power (accuracy
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.
| Thumbnail | Title | Date Uploaded | Visibility | Actions |
|---|---|---|---|---|
|
|
Anakweze_gwu_0075A_17440.pdf | 2025-12-11 | Open Access |
|