Unexpectedly Difficult Statistical Concepts Archives

[123] Dear Political Scientists: The binning estimator violates ceteris paribus

Posted on March 5, 2025March 5, 2025 by Uri Simonsohn

This post delves into a disagreement I have with three prominent political scientists, Jens Hainmueller, Jonathan Mummolo, and Yiqing Xu (HMX), on a fundamental methodological question: how to analyze interactions in observational data? In 2019, HMX proposed the "binning estimator" for studying interactions, a technique that is now commonly used by political scientists. I argued…

[121] Dear Political Scientists: Don't Bin, GAM Instead

Posted on December 3, 2024March 5, 2025 by Uri Simonsohn

There is a 2019 paper, in the journal Political Analysis (htm), with over 1000 Google cites, titled "How Much Should We Trust Estimates from Multiplicative Interaction Models? Simple Tools to Improve Empirical Practice". The paper is not just widely cited, but is also actually influential. Most political science papers estimating interactions now-a-days, seem to…

[120] Off-Label Smirnov: How Many Subjects Show an Effect in Between-Subjects Experiments?

Posted on September 16, 2024September 16, 2024 by Uri Simonsohn

There is a classic statistical test known as the Kolmogorov-Smirnov (KS) test (Wikipedia). This post is about an off-label use of the KS-test that I don’t think people know about (not even Kolmogorov or Smirnov), and which seems useful for experimentalists in behavioral science and beyond (most useful, I think, for clinical trials and field…

[103] Mediation Analysis is Counterintuitively Invalid

Posted on September 26, 2022September 6, 2023 by Uri Simonsohn

Mediation analysis is very common in behavioral science despite suffering from many invalidating shortcomings. While most of the shortcomings are intuitive [1], this post focuses on a counterintuitive one. It is one of those quirky statistical things that can be fun to think about, so it would merit a blog post even if it were…

[99] Hyping Fisher: The Most Cited 2019 QJE Paper Relied on an Outdated Stata Default to Conclude Regression p-values Are Inadequate

Posted on October 13, 2021October 27, 2021 by Uri Simonsohn

The paper titled "Channeling Fisher: Randomization Tests and the Statistical Insignificance of Seemingly Significant Experimental Results" (.htm) is currently the most cited 2019 article in the Quarterly Journal of Economics (372 Google cites). It delivers bad news to economists running experiments: their p-values are wrong. To get correct p-values, the article explains, they need to…

[91] p-hacking fast and slow: Evaluating a forthcoming AER paper deeming some econ literatures less trustworthy

Posted on September 15, 2020August 16, 2021 by Uri Simonsohn

The authors of a forthcoming AER article (.pdf), "Methods Matter: P-Hacking and Publication Bias in Causal Analysis in Economics", painstakingly harvested thousands of test results from 25 economics journals to answer an interesting question: Are studies that use some research designs more trustworthy than others? In this post I will explain why I think their…

[88] The Hot-Hand Artifact for Dummies & Behavioral Scientists

Posted on May 27, 2020November 18, 2020 by Uri Simonsohn

A friend recently asked for my take on the Miller and Sanjurjo's (2018; pdf) debunking of the hot hand fallacy. In that paper, the authors provide a brilliant and surprising observation missed by hundreds of people who had thought about the issue before, including the classic Gilovich, Vallone, & Tverksy (1985 .htm). In this post:…

[80] Interaction Effects Need Interaction Controls

Posted on November 20, 2019February 11, 2020 by Uri Simonsohn

In a recent referee report I argued something I have argued in several reports before: if the effect of interest in a regression is an interaction, the control variables addressing possible confounds should be interactions as well. In this post I explain that argument using as a working example a 2011 QJE paper (.htm) that…

[79] Experimentation Aversion: Reconciling the Evidence

Posted on November 7, 2019February 11, 2020 by Berkeley Dietvorst, Rob Mislavsky, and Uri Simonsohn

A PNAS paper (.htm) proposed that people object “to experiments that compare two unobjectionable policies” (their title). In our own work (.htm), we arrive at the opposite conclusion: people “don’t dislike a corporate experiment more than they dislike its worst condition” (our title). In a forthcoming PNAS letter, we identified a problem with the statistical…

[71] The (Surprising?) Shape of the File Drawer

Posted on April 30, 2018January 23, 2019 by Leif Nelson

Let’s start with a question so familiar that you will have answered it before the sentence is even completed: How many studies will a researcher need to run before finding a significant (p<.05) result? (If she is studying a non-existent effect and if she is not p-hacking.) Depending on your sophistication, wariness about being asked…