Randomized Response Methods for Privacy Protection in Data Collection and Identification Risk Control in Data Release
Open AccessWe develop data masking methods for protecting respondents’ privacy in data collection and preserving data confidentiality when releasing data. We consider only categorical variables, where any mechanism for masking the true values can be viewed as a randomized response (RR) procedure. However, for confidentiality protection, it is called a post-randomization method (PRAM) because the data are randomized after collection. Nonetheless, the problem of devising suitable data masking mechanisms reduces to designing appropriate RR methods. Naturally, designing an RR procedure for privacy protection depends crucially on the privacy criterion. We examine some existing privacy criteria and describe their drawbacks. We show that a previous notion of average security is severely inadequate. Several other criteria, which impose upper bounds on the parity of the RR design, can be satisfied only with substantial data utility loss, unless the number of categories is fairly small. This applies to local differential privacy (LDP), which is a leading privacy criterion. It also shows that the RAPPOR procedure, which has been in use by Google, Apple and others, is highly inefficient. We propose a new privacy procedure that is similar to l-diversity but, works locally for each respondent. The procedure is simple to implement, and its privacy protection is easy to understand and communicate. We give an unbiased estimator of the probability vector of all categories and prove its minimaxity within a class of estimators under squared error loss. We explain that the new procedure offers a better privacy-utility trade-off than LDP.One major concern in releasing microdata is the possibility of identifying the records of some of the units by matching the values of some of the variables, called key variables, which can be obtained easily from other sources. An instance of correctly identifying a unit by matching key variables is called an identity disclosure. It is considered as one of the most serious confidentiality breaches, as it reveals all information for the identified unit. Thus, limiting identity disclosure risks is an important disclosure control goal. A recently developed method, called inverse frequency rule post-randomization, implements upper bounds on the identity disclosure risk. Specifically, for any given ξ > 1/3, it guarantees that the probability of correctly identifying any unit is no more than ξ. However, it cannot give that guarantee for ξ ≤ 1/3. We develop a new unbiased PRAM procedure that can give a similar guarantee even when ξ is much smaller than 1/3. Theoretically, in our case, ξ can be as small as the reciprocal of the sample size. We apply the procedure to a real data set for illustration and an empirical evaluation. We also study the variance inflation due to our PRAM theoretically. We find that the variance inflation is proportional to sampling variance and the proportionality constant increases to 1 as ξ decreases to zero.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.