Securing Training Data in Large Language Models
Open Access DepositedA Synergistic Implementation of Differential Privacy and Homomorphic Encryption to Protect Generative Artificial Intelligence Fine-Tuning Training Data
Generative Artificial Intelligence (GenAI) has witnessed phenomenal growth and is being leveraged by businesses and firms across various sectors. However, existing centralized fine-tuning practices pose inherent problems when dealing with sensitive data residing across different organizations. Medical institutions can't share patient records, financial organizations can't share transaction data, and legal firms can't share client information, yet all could benefit from collaborative model improvement.This research presents a novel approach to privacy-preserving federated fine-tuning that enables multiple organizations to collaboratively improve generative AI models without sharing their private data. We introduce Privacy-Preserving Federated Fine-Tuning (PPFT), a system that uniquely combines three complementary technologies
(1) Parameter-Efficient Fine-Tuning (PEFT) via Low-Rank Adaptation (LoRA), (2) Differential Privacy (DP) for mathematical privacy guarantees, and (3) Homomorphic Encryption (HE) for cryptographic protection during aggregation. Our key insight is that by leveraging LoRA's parameter efficiency, reducing trainable parameters by 99.985%, we can make both differential privacy and homomorphic encryption computationally feasible for neural network training, something previously considered impractical for large-scale models. Through our proof-of-concept implementation using DistilGPT-2, we demonstrate the first successful integration of homomorphic encryption and differential privacy for neural network parameter aggregation in a federated learning setting. Our system achieves formal privacy guarantees of (4.0, 1e-5)-differential privacy and 128-bit cryptographic security while enabling collaborative fine-tuning across organizational boundaries. However, our experiments also reveal significant challenges
the model experienced a 130,000× increase in perplexity compared to non-private baselines, and training time increased by 167% due to cryptographic operations. These results uncover what we term the "Parameter Efficiency Paradox" – while parameter reduction is essential for making privacy-preserving techniques computationally feasible, it simultaneously reduces the model's capacity to learn from noisy, privacy-protected data. Our analysis suggests that achieving practical privacy-preserving federated fine-tuning requires carefully balancing three competing objectives
privacy protection, model utility, and computational efficiency. Despite the challenges, our work establishes that privacy-preserving federated fine-tuning is technically feasible and opens new possibilities for collaborative AI development in privacy-sensitive domains. We provide detailed analysis of the privacy-utility tradeoffs, scalability considerations, and practical deployment strategies. Our findings indicate that with larger datasets (100-1000× more than our experimental setup), optimized implementations, and adaptive privacy mechanisms, this approach could enable practical applications in healthcare, finance, legal services, and other domains where data privacy is paramount. This research contributes to the growing body of work on trustworthy AI by demonstrating that strong privacy guarantees and collaborative learning, while in tension, need not be mutually exclusive. We conclude with a comprehensive discussion of future research directions, including advanced composition techniques, noise-resilient training methods, and privacy-aware neural architectures that could address the current limitations and make privacy-preserving federated learning practical for real-world deployments.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.
| Thumbnail | Title | Date Uploaded | Visibility | Actions |
|---|---|---|---|---|
|
|
Tchankwe_gwu_0075A_17519.pdf | 2025-12-11 | Open Access |
|