Checking your A/B test results every day can inflate your false-positive rate from a nominal 5% to as high as 30% (Optimizely). That's not a typo. If you're like most practitioners, you've probably peeked at an experiment mid-run, felt that surge of excitement when the variation hits significance, and stopped the test early. This is the most common and most damaging mistake in A/B testing.
The Seduction of Early Significance
We've all been there. You launch a test on Monday, and by Wednesday the p-value is 0.04. The variation is winning! Your HiPPO (Highest Paid Person's Opinion) is already drafting the launch email. But here's the problem: with a fixed-sample test, every peek at the data increases your chance of a false positive. Evan Miller, a well-known statistician, showed that if you peek ten times during an experiment, what you think is 1% significance is actually just 5% significance (Evan Miller). In a worst-case scenario, stopping the moment 5% significance appears—after checking after every single observation—can push the actual false-positive rate to 26.1% (Evan Miller). That means more than one in four of your "winners" might be pure noise.
The Only Antidote: Pre-Commit or Go Sequential
The fix is straightforward: decide your sample size before you start, and don't stop early. If you must peek, use a sequential testing method that adjusts for multiple looks. This isn't just academic theory—it's how you avoid shipping bad changes. At Bing, where over 250 experiments run daily and almost every user is in about 15 simultaneous experiments (Bing), false positives are a constant threat. They've built systems to alert experimenters to statistical interactions between experiments (Kohavi et al.). If a company with billions of users and a dedicated stats team still worries about false positives, you should too.
But What About Bayesian? Isn't It Better?
Some argue that Bayesian methods let you peek all you want without inflating false positives. That's true—Bayesian testing allows continuous monitoring and outputs probabilities like "92% chance B is better" (Optimizely, VWO). But here's the catch: Bayesian methods often require prior knowledge, and if you set a weak prior or misinterpret the posterior, you can still be fooled. More importantly, switching to Bayesian doesn't absolve you from the discipline of pre-registering your hypotheses and metrics. The real problem isn't frequentist vs. Bayesian—it's peeking without a plan. I've seen teams adopt Bayesian tools and still stop early because the probability hit 95%. The tool doesn't fix the behavior.
What the Data Really Shows
Let's look at the evidence. When Optimizely re-analyzed 48,000 historical experiments, they found that fixed-horizon statistics declared a winner or loser in 36% of tests, while their sequential Stats Engine (which allows continuous monitoring) declared a winner or loser in only 22%—uncovering 39% fewer conclusive results (Optimizely Stats Engine). That's not because sequential testing is underpowered; it's because many of those fixed-horizon "winners" were false positives. In another analysis, they found that using classical statistics with peeking and multiple metrics increased the chance of incorrectly declaring a winner by over 5x (Optimizely Stats Engine). The cost is real: at Microsoft, a 1% revenue improvement is worth about $10 million annually in the US (Kohavi et al.). Shipping a false positive that actually hurts revenue by 1% is a $10 million mistake.
Practical Steps to Save Your Tests
Here's what I'd actually do. Before you launch any test, calculate the sample size you need based on your baseline conversion rate and the minimum effect you care about. Use a calculator like Evan Miller's, or the rule of thumb: n = 16 * (variance / delta^2) (Evan Miller). Decide on your primary metric and stick to it. If you're tempted to peek, set up a sequential test that gives you always-valid p-values—tools like Optimizely's Stats Engine or the SPRT (Sequential Probability Ratio Test) can reduce required sample sizes by up to 50% or more (Evan Miller, Optimizely). And if you're running A/A tests, do them quarterly to verify your setup isn't broken (Optimizely). Remember, an A/A test that shows significance at 95% is a red flag (Optimizely).
Don't let the excitement of a potential win blind you to the statistics. Pre-commit, use sequential methods, and validate your pipeline. Your future self—and your company's revenue—will thank you.
Sources
- Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
- Evan Miller - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
- Bing - https://blogs.bing.com/search-quality-insights/August-2013/Large-Scale-Experimentation-at-Bing/
- Kohavi et al. - https://exp-platform.com/large-scale/
- Optimizely A/A testing glossary - https://www.optimizely.com/optimization-glossary/aa-testing/
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!