Skip to main content
Optimization Tools

Stop Peeking: How to Run A/B Tests Without Fooling Yourself

Peeking at A/B test results can inflate false positives to 30%. Learn how to pre-commit to sample sizes, use sequential testing, and trust the process.

If you check your A/B test results every day and stop as soon as you see 95% significance, you're fooling yourself. That daily peeking can inflate your false-positive rate from a nominal 5% to as high as 30% (Optimizely). In other words, nearly one in three of your "winners" could be pure noise.

This is for anyone who runs A/B tests — marketers, product managers, UX designers — and wants to make decisions that actually stick. I'm going to walk you through a practical process that keeps you honest, from pre-committing to a sample size to knowing when to break the rules.

1. Pre-commit to a sample size before you start

The single biggest mistake I see is people deciding when to stop an experiment based on how the data looks. That's backwards. You should decide how many visitors you need before you launch. A classic rule of thumb is n = 16 times sigma-squared divided by delta-squared, where delta is the minimum effect you care about and sigma-squared is the variance (Evan Miller). For a conversion rate, variance is p(1-p), so if your baseline is 10%, sigma-squared is 0.09.

Let's make this concrete. Suppose your current landing page converts at 10%, and you'd be excited to detect a 20% relative improvement (so delta = 0.02, going from 10% to 12%). Plug in: n = 16 * 0.09 / (0.02^2) = 16 * 0.09 / 0.0004 = 3,600 visitors per variation. That's your number. You need at least 3,600 visitors in each group to have a decent shot at detecting that effect.

Why bother? Because if you decide on the sample size in advance, you're far less tempted to peek and stop early. The math only works if you stick to the plan.

2. Use a fixed end date or sequential testing

Once you've set your sample size, the natural question is: how long will it take? Divide the total visitors you need by your average daily traffic. If you need 7,200 total visitors and you get 1,000 per day, that's about a week. Fine. But here's the catch: you can't just check the data every day and decide to stop early when it looks good. That's the peeking problem.

What can go wrong? If you peek ten times, what you think is 1% significance is actually just 5% significance (Evan Miller). And if you stop the moment you see 5% significance, your actual false-positive rate can hit 26.1% in the worst case (Evan Miller). That's a warning you should take seriously.

The defense is simple: either commit to a fixed end date or use a sequential testing method that adjusts the thresholds for multiple looks. Optimizely's Stats Engine, for example, lets you peek anytime and still get valid results, and it can call a winner up to 2.5 times faster for experiments over 50,000 visitors (Optimizely).

3. Watch out for practical significance, not just p-values

Statistical significance tells you a result isn't due to chance, but it says nothing about whether that result matters. With a huge sample, even a microscopic improvement can become statistically significant. A 0.2 percentage point difference in completion rate (85.0% vs 85.2%) can be statistically significant (p = 0.03) yet be too small to justify any design change (Nielsen Norman Group).

So before you celebrate a p-value of 0.001, ask yourself: is the effect size large enough to care about? The minimum detectable effect you chose in step 1 is your guide. If the observed lift is smaller than that, it's probably not worth shipping.

Here's a real-world scale: a 1% improvement to revenue at Microsoft is worth about $10 million annually in the US (Kohavi et al.). That's a hefty prize. But if your test only moves a metric by 0.1%, maybe it's not worth the engineering time.

4. Consider A/A tests and sample ratio mismatch

Even with a perfect plan, your experiment setup can betray you. That's where A/A tests come in. Run two identical versions against each other, and if you see a statistically significant difference (greater than 95% significance) between the two, something is wrong with your implementation (Optimizely A/A testing glossary). It's a good idea to run these calibration checks quarterly (Optimizely A/A testing glossary).

Another red flag is sample ratio mismatch (SRM), where the number of visitors in each group doesn't match your intended ratio. SRM is like a fever — it's a symptom of various underlying issues (KDD 2019). If you ignore it, you might ship a bad product modification that appeared to be good, or vice versa (KDD 2019). So if you see SRM, stop and investigate before drawing any conclusions.

Quick tip: if your A/A test flags a difference, don't just re-run it — check your tracking code, your randomization logic, and your data pipeline.

What can go wrong? The most common failure I see is someone running an A/B test, seeing a 95% significance after two days, and stopping. They then ship a change that actually has no effect — or worse, a negative effect — because they didn't pre-commit and peeked their way into a false positive.

The most important thing to remember: Decide your sample size before you start, stick to it, and let the test run its course. Your future self will thank you.

Sources

  • Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
  • Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
  • Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
  • Nielsen Norman Group - https://www.nngroup.com/articles/practical-significance/
  • Optimizely A/A testing glossary - https://www.optimizely.com/optimization-glossary/aa-testing/
  • KDD 2019 - Diagnosing Sample Ratio Mismatch - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom

Share this article:

Comments (0)

No comments yet. Be the first to comment!