Machine Learning for Non-Invasive Cyber Insurance Adversarial Underwriting
Open AccessIn light of the growing dependence on digital technology, organizations are prone to cyberattacks that cannot be entirely mitigated, hence precluding the attainment of absolute or 100% security. In light of the growing complexity and prevalence of cyber risks, many businesses are progressively resorting to cyber insurance as a means to mitigate the financial ramifications resulting from a successful cyberattack. Nonetheless, the inadequate evaluation of policyholders' cyber risk exposure and inaccurate underwriting practices result in the leakage of cyber insurance premiums and pose a solvency risk that may compel insurance providers to declare bankruptcy or to completely exit the cyber insurance market. The legal disputes and subsequent claims settlements between Mondelez and Zurich Insurance Group ($100 million) or between Merck and Ace American Insurance ($1.4 billion) after the 2017 NotPetya ransomware outbreak serve as illustrative instances that highlight the significance of this subject matter.A key metric for measuring the profitability of a line of insurance is the statutory Direct Loss plus Defense and Cost Containment Expenses (DCCE) ratio, which represents the portion of the annual earned premium (EP) spent on claims. As a result of faulty underwriting decisions and the growing rate of cyberattacks, the direct loss ratio for standalone cyber insurance in the global market has been steadily rising and is above 60% since 2020. It has become essential to leverage adversarial underwriting to establish a quantitative framework for improving the efficiency and effectiveness of the cyber insurance industry. Adversarial underwriting is all about predicting cyber risk using attackers’ tactics, techniques, and procedures (TTPs). Data scientists can accurately evaluate and predict cyber risk by understanding the logical patterns and behaviors of hackers. Therefore, our study investigated the use of centralized corporate security incident events logs (SIEM) by machine learning models to anticipate policyholders’ cybersecurity exposure. The aim was to reduce cyber underwriting risks for insurance providers and provide equitable cyber coverage premiums for policyholders. Nevertheless, the unavailability of information such as SIEM logs in the public domain is attributed to potential cybersecurity risks and privacy concerns. In fact, SIEM logs include consolidated network event alerts that have the potential to reveal confidential information and undermine a company's reputation. Consequently, CIC-DDoS2019, DAPT2020, and UNSW_NB15 cybersecurity datasets were used to test our prototype and achieve our research goals. We fed the three datasets into the ELK (Elasticsearch Logstash Kibana) Stack SIEM to ensure that the signals or features learned by the proposed deep learning models during training are present in the log data that will feed the model during production or inference. In order to develop a collection of predictive models for the identification of specific cyber threats, we used machine learning frameworks such as TensorFlow and Keras on Google Colaboratory. Additionally, our study leveraged Weights & Biases for the optimization of model’s hyperparameters, Hopsworks for managing the feature store and storing trained models, and KServe on Kubernetes (K8s) for serving the models. Our research results are expected to provide a valuable contribution to the current body of literature on the insurability and underwriting of cyber risks.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.