You've just launched a new landing page, and by the afternoon you're already refreshing the dashboard. The variation is up 12%, and the p-value is under 0.05. You want to call it a win. But here's the blunt truth: that number is probably a lie. Peeking at your results and stopping early is the single most common way to break an A/B test, and it can inflate your false-positive rate from a nominal 5% to as high as 30% (Optimizely A/B testing). If you're running tests without a pre-committed sample size, you're not running experiments—you're gambling.
The Peeking Trap: Why Your Results Are Untrustworthy
Picture this: you're the product manager at a mid-sized SaaS company. You've just redesigned your pricing page, and you want to know if the new layout boosts sign-ups. You set up a test with a control and a variation, and you start checking the results every morning. By day three, the variation shows a 6% lift with a p-value of 0.04. Your instinct says, "Ship it!" But that instinct is exactly what Evan Miller warned about when he wrote "How Not to Run an A/B Test." He showed that if you peek at an ongoing experiment ten times, what you think is 1% significance is actually just 5% significance (Evan Miller - How Not to Run an A/B Test). In a worst-case scenario, stopping as soon as 5% significance appears—while running a significance test after every observation—can push the actual false-positive rate to 26.1% (Evan Miller - How Not to Run an A/B Test). That's not a rare edge case; it's the norm for anyone who doesn't plan ahead.
Decide Your Sample Size Before You Launch
So how do you fix this? You decide on a sample size before you even write a line of code. The core question is: what's the smallest improvement you care about? That's your minimum detectable effect (MDE). If your current conversion rate is 5%, and you'd only bother shipping a change that lifts it to 5.5%, your MDE is 10% relative. You also need to pick your significance level and power—usually 0.05 and 80% (Optimizely A/B testing). A rule of thumb from Evan Miller's sample size guide is n = 16 times (sigma-squared / delta-squared), where delta is the MDE and sigma-squared is the variance of your metric (Evan Miller - How Not to Run an A/B Test). For a conversion rate, that variance is p(1-p). So if p = 0.05, sigma-squared = 0.0475. If delta = 0.005 (a 0.5 percentage point lift), you'd need n = 16 * (0.0475 / 0.000025) = 30,400 visitors per variation. That's a lot, but it's the honest price of detecting a small effect.
Why Small Effects Are (Usually) Not Worth It
Now, here's where you need to get real about what matters. The Nielsen Norman Group makes a crucial distinction between statistical significance and practical significance (Nielsen Norman Group - Practical Significance). A result can be statistically significant and still be useless. For example, a completion-rate difference of 0.2 percentage points—85.0% versus 85.2%—can show p = 0.03 with a huge sample, but that difference is too small to justify a design change (Nielsen Norman Group - Practical Significance). At Microsoft, they've learned that a 1% improvement to revenue is worth about $10 million annually in the US (Kohavi et al. - Online Controlled Experiments at Large Scale (KDD 2013)). That sounds big, but many ideas that would move a key metric by 1% are not well estimated in advance. So when you're setting your MDE, think about what the business actually needs—not what would make a pretty chart.
Sequential Testing: The Smarter Way to Peek
But let's be honest—you're going to peek. The human brain can't resist looking at a live experiment. That's why sequential testing exists. Instead of trying to ignore the data, you embrace it with methods that adjust your thresholds for multiple looks. Evan Miller's simple sequential procedure is a great starting point: you choose a sample size N, assign users 50/50, and stop when the treatment-minus-control success count reaches 2 times the square root of N (declaring a winner) or the total success count reaches N (declaring no winner) (Evan Miller - Simple Sequential A/B Testing). This method can reduce the number of observations needed for a successful experiment by 50% or more, and it works extremely well with low conversion rates (Evan Miller - Simple Sequential A/B Testing). The bigger idea is always-valid inference, which lets you look at results any time and still get valid answers (arXiv - Always Valid Inference (Johari et al.)). That's what Optimizely's Stats Engine does—it was built with Stanford statisticians and can call a winner or loser up to 2.5 times as fast as fixed-horizon statistics for experiments over 50,000 visitors (Optimizely Stats Engine). The catch: sequential methods can be more conservative, so you might not get a conclusive result as often. In fact, when Optimizely re-ran 48,000 historical experiments, fixed-horizon statistics declared a winner or loser in 36% of tests, while the sequential Stats Engine did so in only 22%—that's 39% fewer conclusive results (Optimizely Stats Engine). So you trade speed for certainty.
Use Pre-Experiment Data to Cut Your Sample Size
Another trick to make your test faster is CUPED, which stands for Controlled-experiment Using Pre-Experiment Data. It uses data from before the experiment to adjust your metrics and reduce variability (Deng et al. - CUPED (WSDM 2013)). At Bing, three recent experiments using CUPED showed variance reductions of 45%, 52%, and 49% with just one week of pre-experiment data (Deng et al. - CUPED (WSDM 2013)). Cutting variance by about 50% effectively means you need only half the users or half the duration to get the same statistical power (Deng et al. - CUPED (WSDM 2013)). So if your sample size calculation says you need 30,000 visitors, CUPED might get you down to 15,000. That's a huge win for a busy product team.
Check Your Test Setup: A/A Tests and Sample Ratio Mismatch
Finally, before you trust any result, make sure your test is actually working. Run an A/A test—two identical versions—to validate your setup. Optimizely recommends doing this quarterly (Optimizely A/A testing glossary). If an A/A test reports a statistically significant difference greater than 95% between two identical versions, that's a red flag that your implementation is broken (Optimizely A/A testing glossary). Also watch for a Sample Ratio Mismatch (SRM), where the observed sample ratio differs from the expected one (KDD 2019 - Diagnosing Sample Ratio Mismatch). An SRM is like a fever—it's a symptom of a variety of data quality issues (KDD 2019 - Diagnosing Sample Ratio Mismatch). If you ignore it, you might ship a bad product modification that looks good, or vice versa (KDD 2019 - Diagnosing Sample Ratio Mismatch). So before you celebrate that 12% lift, run an A/A test and check your sample ratios.
Quick tip: If you can't resist peeking, use a sequential test or a fixed end date—and never stop early just because the p-value dips below 0.05.
Sources
- Optimizely - A/B testing: https://www.optimizely.com/optimization-glossary/ab-testing/
- Evan Miller - How Not to Run an A/B Test: https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Optimizely - Stats Engine: https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
- Evan Miller - Simple Sequential A/B Testing: https://www.evanmiller.org/sequential-ab-testing.html
- Deng et al. - CUPED: https://exp-platform.com/Documents/2013-02-CUPED-ImprovingSensitivityOfControlledExperiments.pdf
- KDD 2019 - Diagnosing Sample Ratio Mismatch: https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!