Anonymization and De-Anonymization Attacks in Online Social Networks
Open AccessOnline social networks have gained tremendous popularity recently. Millions of people use social network apps to share precious moments with friends and family. Users are often asked to provide personal information such as name, gender, address when using social networks. This information could be collected, analyzed, and re-published at large scale for both academic and business studies, and are usually processed by anonymization methods to protect user's privacy. However, these anonymized data could be misused by unauthorized third parties and even attackers to violate users' privacy. Several structure-based de-anonymization techniques have been proposed to re-identify the users in anonymized networks. In light of this, we first address the novel problems of re-identifying users in anonymized social networks with the help of public user attributes in this work. More specifically, we quantify the significance of attributes in a social network, based on which we propose an attribute-based similarity measure; then we design an algorithm by exploiting attribute-based similarity to de-anonymize social network data; finally we employ the dataset collected from a real-world online social network to evaluate our method. And experimental results show that public user attributes can significantly improve the de-anonymization accuracy.Extensive researches on anonymization techniques have been carried out to protect the data from privacy violations in social networks. Nevertheless, anonymization may affect the usability of the data as random noises are introduced, decreasing user experience. Therefore, a trade-off between privacy protection and data usability must be sought. In this work, we employ various graph and application utility metrics to investigate this trade-off. More specifically, we conduct an empirical study by implementing five state-of-the-art anonymization algorithms to analyze the graph and application utilities on a Facebook and a Twitter dataset. Our results indicate that most anonymization algorithms can partially or conditionally preserve the graph and application utilities and any single anonymization algorithm may not always perform well on different datasets.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.