Electronic Thesis/Dissertation
 

Identification Risk Control in Microdata Release by Inverse Frequency Post-Randomization and Its Impact on Data Utility

Open Access

e a masked or perturbed version of survey data to protect respondents' confidentiality. Ideally, a perturbation procedure should protect confidentiality without much loss of data quality, so that released data may practically be treated as original data for making inferences. A major objective in releasing microdata is to control the risk of correctly identifying any respondent's records, by matching the values of some identifying or key variables. For categorical key variables, we propose a new approach to measuring identification risk and setting strict disclosure control goals. The general idea is to ensure that the probability of correctly identifying any respondent or surveyed unit is at most ξ, which is pre-specified. Then, we develop an unbiased post-randomization procedure that achieves this goal for ξ >1/3. The procedure allows substantial control over possible changes to the original data and the variance it induces is of a lower order of magnitude than sampling variance under multinomial sampling scheme. We apply the procedure to a real data set, where it performs consistently with the theoretical results and quite importantly, shows very little data quality loss. We also propose a variation of our procedure that achieves little data utility loss for estimating population totals under general probability sampling schemes.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Zhang_gwu_0075A_14124.pdf Zhang_gwu_0075A_14124.pdf 2018-05-02 Open Access