A Framework for Modeling Data Breach Risk Using Machine Learning Models for High-Dimensional Panel Data
Open AccessA cybersecurity data breach has become one of the crucial issues for organizations. While facing immediate financial impact, organizations lose their customers' faith as data breaches result in the theft of consumers' personally identifiable data. Startup business ventures are collapsing as an increasing number of data breach incidents lead to the theft of their intellectual property before producing or releasing a product on the market. Organizations have the responsibility to release these breach incidents to the public. This information on individual data breaches helps identify risk factors associated with probability and size of data breaches, duration until the subsequent data breach, analysis of geographic differences, and data breach features that are trending over time. Such comprehensive analysis would help organizations better manage data breach risks and gain executive management support in implementing security systems, including but not limited to intrusion detection systems. This research introduces a framework for analyzing publicly available panel data for risk factor identification and early warning predictions of data breach risks. We develop a methodology for quantitative analysis of publicly available data breach records aimed at rigorous identification of trending data breach characteristics, sources of geographical heterogeneity of data breach incidence in the US, as well as on estimation of probability, size, and timing of data breaches in a given organization. We developed a series of supervised machine learning models that predict the probability of data breach incidence, size, and timing for any given US organization. By analyzing feature importance, partial dependence, and hazard ratios, this research also revealed early warning signals of data breach incidence, size, and timing for US organizations. The proposed modeling framework is based on tree-based supervised machine learning methods adapted to high-dimensional sparse panel data and on a set of nonparametric and parametric survival analysis techniques which have been applied to data breach modeling for the first time and have been shown to provide a promising toolbox that allows directly addressing the question about the timing of repeat data breaches. This study's results can help security engineers assess their organization's susceptibility to data breach risks based on various contextual features and developers of data security systems to determine which regions, industries, and types of breaches, to target.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.