All is Fair in Chat and Code: Leveraging GenAI Techniques to Detect Advanced Threats of Malware and IP Leakage in LLM-enabled Web Applications
Open Access DepositedIntroduced in January 2024, The OpenAI GPT Store is a dynamic marketplace for custom versions of ChatGPT. The GPT Store hosts over three million GPTs that were created by more than a million AI developers. The diverse nature and significant utility of the GPT Store illustrates the transformative potential of LLMs towards becoming general-purpose technologies that are capable of driving innovation across industries. Despite being in its formative stages, the GPT Store commands over 1.5% of ChatGPT’s total desktop visits, indicating a rapidly growing, dedicated user base, which is primarily concentrated around a select number of popular GPTs. Established literature has already ascertained that the top 1% and top 5% of apps in the store are responsible for 90% and 95% of daily conversations respectively on the platform. However, GPTs exhibit structural vulnerabilities that can be exploited by malicious actors. These vulnerabilities stem from their instruction-following nature, allowing attackers to manipulate inputs to invoke unintended actions, such as accessing confidential information (Intellectual Property Leakage) or executing malware. Indirect prompt injection attacks represent a significant risk, echoing longstanding security challenges, and leading to GPTs being likened as the “new” AI spin on the age-old application security problem of malicious input and data leakage. This praxis proposes a fundamental change in the manner in which indirect prompt injection vulnerabilities on the GPT marketplace have been assessed by the research community to date; shifting from evaluating vulnerabilities from a product existence perspective, towards a more comprehensive view that encompasses prevalence (usage patterns), discoverability and credibility. Through this paradigm shift, we are able to quantify and showcase how the severity of risk on the platform has been under-stated to date. Employing AI classification techniques, we determine GPTs that prompt users to take actions outside of the OpenAI ChatGPT platform and which contain malware, spam, phishing or criminal intellectual property, are 30.8% more utilized than benign GPTs. When accounting for all conversations associated with GPTs that contain malicious, suspicious or unrated domain URLs, we observe 1 in 20 conversations that take place on the OpenAI Store, as being at risk for phishing or malware attacks. Employing Generative AI techniques, we identify GPTs that are susceptible to intellectual property (IP) leakage, and establish a 22.15% IP leakage rate (approximately 1 in every 5 GPTs) using our constructed praxis dataset (which represents over 95% of daily conversations). Alarmingly, over 1 in 3 GPTs listed on the OpenAI Storefront (which showcases top performing GPTs in each category) leak information against the expressed intent of the AI Systems designer. We use a series of ROUGE evaluation metrics to evaluate the quality of our GenerativeAI pipelines, establishing an industry baseline score of 0.81. This score suggests that our proposed prompts, parameters, and GenAI pipelines perform well as a baseline.We further combine these AI/ML and GenAI threat detection techniques into a novel, comprehensive GPT vulnerability assessment tool, aptly named SALUS Detector. Named after the Greek God of Safety, SALUS takes as an input a GPT g-id (global identifier associated with each GPT listed on the store, and which can be extracted from the GPT url), and returns a vulnerability score, with flags denoting key areas of risk, both from a user (demand-side) and developer’s (supply-side) perspective. Our novel tool serves as a “GPT anti-virus scanner” for users, and a “GPT compliance and security companion” for AI systems designers to understand the GPT’s security profile. This praxis’s novel contributions are twofold: (1) We evolve the threat detection landscape based on a new, evaluation grain (usage versus product count), and employ the use of intent analyses on unstructured assistant responses; (2) We propose a novel suite of analysis techniques, leveraging GenerativeAI technologies to detect and mitigate against the cacophony of technological, operational, economic, and societal risks that indirect prompt injection attacks promulgate. We package these learnings into a new tool, named SALUS, which aims to enhance the resilience and security of the GPT ecosystem. SALUS seeks to ensure that as the underlying LLMs that underpin GPTs expand in both their capability and reach, that users and developers are able to safeguard themselves against potential misuse and/or emerging threats. We would like to submit SALUS for a patent.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.