首页 | 本学科首页   官方微博 | 高级检索  
     检索      


Two-sample tests for sparse high-dimensional binary data
Authors:Amanda Plunkett
Institution:Department of Defense, Fort Meade, MD, USA
Abstract:In this article, we study the methods for two-sample hypothesis testing of high-dimensional data coming from a multivariate binary distribution. We test the random projection method and apply an Edgeworth expansion for improvement. Additionally, we propose new statistics which are especially useful for sparse data. We compare the performance of these tests in various scenarios through simulations run in a parallel computing environment. Additionally, we apply these tests to the 20 Newsgroup data showing that our proposed tests have considerably higher power than the others for differentiating groups of news articles with different topics.
Keywords:Dimension reduction  High dimension  Multivariate binary data
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号