首页 | 本学科首页   官方微博 | 高级检索  
     


Post-randomization for controlling identification risk in releasing microdata from general surveys
Authors:Cheng Zhang  Tapan K. Nayak
Affiliation:aMedStar Cardiovascular Research Network, Washington, DC, USA;bCenter for Statistical Research and Methodology, U.S. Census Bureau, Washington, DC, USA;cDepartment of Statistics, George Washington University, Washington, DC, USA
Abstract:
Before releasing survey data, statistical agencies usually perturb the original data to keep each survey unit''s information confidential. One significant concern in releasing survey microdata is identity disclosure, which occurs when an intruder correctly identifies the records of a survey unit by matching the values of some key (or pseudo-identifying) variables. We examine a recently developed post-randomization method for a strict control of identification risks in releasing survey microdata. While that procedure well preserves the observed frequencies and hence statistical estimates in case of simple random sampling, we show that in general surveys, it may induce considerable bias in commonly used survey-weighted estimators. We propose a modified procedure that better preserves weighted estimates. The procedure is illustrated and empirically assessed with an application to a publicly available US Census Bureau data set.
Keywords:Identity disclosure   data partitioning   key variable   post-randomization block   survey weighted estimator   total variation distance
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号