Electronic Thesis/Dissertation
 

Utilizing Agentic AI for Malware Authorship Attribution

Open Access Deposited

As distribution of advanced malware has become more widespread, the need to gather more detailed information about their intent and makeup has become required. One such way of obtaining this information is by understanding the authors of the malware and their motivations. This has traditionally been a labor-intensive process that is becoming increasingly resource constrained due to the sophistication and sheer volume of samples in circulation. Current attempts at automation have been utilized with success, however often lack a modular, explainable and scalable method to coordinate multiple knowledge sources for optimal coverage.This praxis investigates an agentic artificial intelligence (AI) framework designed to automate malware attribution through the integration of agents for structural analysis and behavioral analysis as well as a reasoning component to coordinate multi-agent workflows. The two agents used for static analysis and dynamic analysis, respectively, were each trained using the CatBoost classifier. To conduct reasoning for the agents, as well as provide justification for attribution decisions, three large language models (LLMs) were evaluated, OpenAI’s GPT-4o, GPT-4o-mini, and Anthropic Claude 3.7 Sonnet. A labeled dataset was developed based on the APTMalware repository (Cyber-Research, 2019) and was then enriched with static analysis features and dynamic analysis features created by using the LIEF library and Record Future Triage tool, respectively. This resulted in an agentic AI workflow that achieved over 90% accuracy at a 95% confidence level and demonstrated its efficacy and ability to implement authorship attribution activities. Static analysis agents, which evaluated the structural features of the binaries, consistently outperformed dynamic analysis agents, which looked at the behavioral features. GPT-4o-mini performed comparably to the larger LLMs, showing that lightweight LLMs can provide the same performance as effective reasoning agents at a lower cost and parameter count. These findings support the idea that agentic systems are scalable and viable for malware authorship attribution.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Beck_gwu_0075A_17717.pdf Beck_gwu_0075A_17717.pdf 2025-12-15 Open Access