Advanced Privacy-Preserving Decentralized Federated Learning for Insider Threat Detection in Collaborative Healthcare Institutions
Open Access DepositedThe healthcare sector is embracing a digital transformation driven by technological advancements and the quest for more effective healthcare delivery. This transformation has led to an exponential increase in the volume and diversity of data generated by healthcare institutions. Such large datasets offer the potential to reveal significant insights, including patterns that can predict, detect, and prevent insider threats in collaborative healthcare settings (Brauneck et al., 2023). However, this data growth also intensifies risks related to insider threats, which often originate from within an organization and pose significant challenges due to digital health adoption and privacy concerns with centralized methods, leading to data breaches and increasing regulatory challenges (Alder, 2023; Vasko and Gardner, 2019; Fisher and Phillips, 2024; Teo et al., 2024; U.S. Department of Health and Human Services, 2022). Additionally, the rising adoption of machine learning (ML) for big data analysis and research involves centralized machine learning methods widely applied to tackle these threats. However, these methods face challenges such as privacy violations, ownership disputes, and regulatory non-compliance (Goyal & Malviya, 2023; Tabassum et al., 2024; Paracha et al., 2024; Naresh, Thamarai & Allavarpu, 2023). To mitigate these risks, the centralized federated learning (CFL) method offers improvements through a collaborative approach. However, CFL relies on a main server to aggregate updates from distributed models, which could introduce a single point of failure. This study proposes an Advanced Privacy-Preserving Decentralized Federated Learning (APPDFL) system as a solution for insider threat detection that adopts collaborative healthcare institution participation for intelligence sharing. Leveraging a decentralized federated learning communication architecture, this approach ensures privacy, scalability with large datasets, and flexibility compared to traditional methods (Nguyen et al., 2022). We aim to apply security by design by implementing server-less peer-to-peer (P2P) communication in a horizontal decentralized federated learning with differential privacy mechanism. This mechanism has been widely implemented in numerous domains, such as deep learning research, to prevent indirect data leakage (Abadi et al., 2016; McMahan, 2017). APPDFL allows multiple healthcare institutions to collaboratively train shared threats intelligence without exchanging confidential data, eliminating the need for a central server (Roy et al., 2019; He et al., 2020; LeCun, Bengio, and Hinton, 2015; Shayan et al., 2019). Additionally, federated learning complies with laws and regulations, ensuring data privacy preservation and data security (Saha et al., 2024; Truong et al., 202; Chalamala et al., 2022; Sharma & Kumar, 2023), making it ideal for sensitive sectors such as healthcare, financial institutions, and government agencies. For this Praxis implementation, we utilized the most widely used dataset for insider threat from Carnegie Mellon University-CERT, specifically versions r4.2, r5.2, and r6.2. We applied various balancing techniques to balance the dataset and simulate real-world cases of insider threat detection across multiple healthcare institutions. It is worth noting that to address the challenge presented by literature works that most commonly use only one or two versions of CERT’s dataset and to ensure a fair comparison for validating RH1, traditional, modern, and centralized federated learning approaches were tested on the same dataset versions used for our Advanced Privacy-Preserving Decentralized Federated Learning (APPDFL) system as a proposal solution. The performance of the proposed APPDFL system as a proposed solution is then assessed and compared to both previous works review and centralized methods tested. The trade-off between accuracy and privacy enhancement is evaluated to determine the noise level tolerance that maintains acceptable model accuracy with privacy-preservation among all participants (Ahamed et al., 2024). The proposed APPDFL system as a solution for insider threat detection in collaborative healthcare institutions significantly improves accuracy, enhances privacy, and outperforms both the centralized methods tested and existing work in the field. The findings clearly indicate that collaboration among different participating healthcare institutions for insider threat intelligence sharing is possible. This collaboration maintains superior effectiveness while ensuring the accuracy of the learning process and acceptable noise tolerance. Furthermore, this approach adheres to the core principles of information security: ensuring confidentiality, maintaining integrity, and guaranteeing availability of data—commonly known as the “CIA Triad.”
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.