A Proposed Hybrid Effect Size Plus p-Value Criterion: Empirical Evidence Supporting its Use期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

按检索

A Proposed Hybrid Effect Size Plus p-Value Criterion: Empirical Evidence Supporting its Use

Authors:	William M Goodman Susan E Spruill Eugene Komaroff

Institution:	1. University of Ontario Institute of Technology, Ontario, Canada;2. bill.goodman@uoit.ca;4. Applied Statistics and Consulting, Spruce Pine, NC;5. Graduate School, Keiser University, Fort Lauderdale, FL

Abstract:	ABSTRACT When the editors of Basic and Applied Social Psychology effectively banned the use of null hypothesis significance testing (NHST) from articles published in their journal, it set off a fire-storm of discussions both supporting the decision and defending the utility of NHST in scientific research. At the heart of NHST is the p-value which is the probability of obtaining an effect equal to or more extreme than the one observed in the sample data, given the null hypothesis and other model assumptions. Although this is conceptually different from the probability of the null hypothesis being true, given the sample, p-values nonetheless can provide evidential information, toward making an inference about a parameter. Applying a 10,000-case simulation described in this article, the authors found that p-values’ inferential signals to either reject or not reject a null hypothesis about the mean (α?=?0.05) were consistent for almost 70% of the cases with the parameter’s true location for the sampled-from population. Success increases if a hybrid decision criterion, minimum effect size plus p-value (MESP), is used. Here, rejecting the null also requires the difference of the observed statistic from the exact null to be meaningfully large or practically significant, in the researcher’s judgment and experience. The simulation compares performances of several methods: from p-value and/or effect size-based, to confidence-interval based, under various conditions of true location of the mean, test power, and comparative sizes of the meaningful distance and population variability. For any inference procedure that outputs a binary indicator, like flagging whether a p-value is significant, the output of one single experiment is not sufficient evidence for a definitive conclusion. Yet, if a tool like MESP generates a relatively reliable signal and is used knowledgeably as part of a research process, it can provide useful information.

Keywords:	NHST Minimum effect size plus p-value criterion MESP statistical evidence meaningful distance true power true Type I error rate

设为首页 | 免责声明 | 关于勤云 | 加入收藏