Classification of Deepfake Satellite Images Containing Adversarial Perturbations using Machine Learning
Open Access DepositedAdvances in generative AI have reached the point where humans cannot reliably tell the difference between real and fake images. Methods for both the generation of, as well as detection of deepfakes continue to grow at a rapid pace. While the detection of deepfakes continues to improve, the problem remains unsolved and is projected to have global economic impact and potential national security implications. Additionally, the vulnerability of ML models to adversarial perturbations is a known problem, with some attacks able to achieve 100% attack success rate. While much research has been done on human subjects as the target of deepfakes, this research is based on the novel domain of deepfake satellite imagery, a growing threat with the increased access to training data in this domain. Additionally, most research on deepfake detection has been done on GAN-generated data. This research considers GAN and Diffusion Model (DM) generative techniques. The praxis compares two methods for binary classification of images. The first method is adversarial purification by preprocessing images with the PDM-Pure pixel diffusion model. The second method utilizes the Disjoint Diffusion Deepfake Detection (D4) model. Both methods are evaluated using two different black-box attacks on both GAN-generated as well as DM-generated datasets. The research demonstrates two black-box adversarial attacks are 100% successful with the novel satellite imagery target. Comparison of the detection models results in greatly reduced attack success rate to just 1% in some cases. Both models achieved higher average precision on the GAN-generated data. The D4 model had average precision as high as 98% for GAN-generated satellite deepfakes containing adversarial perturbations, compared with a 90% average precision achieved by the PDM-Pure method under the same conditions. In conclusion, D4 proves to provide a more robust solution for classification of GAN and DM generated deepfake satellite imagery.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.