Skip to main content
Optimization Tools

A/B Tests Keep Lying to You: Stop Peeking and Use Sequential Testing

If you stop your A/B test the moment it hits 95% significance, you're fooling yourself. Learn why peeking is dangerous and how sequential testing fixes it.

You've got a new landing page, and you're checking the results every morning. On day three, the variation hits 95% significance, and you pull the trigger. Sound familiar? Here's the blunt truth: that win might be a mirage. The more you peek, the more likely you are to declare a winner that doesn't exist. In fact, if you check your test daily and stop as soon as significance appears, your false-positive rate can balloon from 5% to as high as 30% (Optimizely). That's not a rounding error; that's a broken decision engine.

Why Peeking Is a Silent Killer

In classical A/B testing, the math assumes you pick a sample size, run the test, and look at the data exactly once. The p-value you see at the end is valid only under that single-look assumption. But if you peek every day and stop early, you're essentially running multiple significance tests on the same data. Every look increases your chance of a false positive. Evan Miller, who has written extensively on this, puts it starkly: if you peek ten times, what you think is 1% significance is actually just 5% significance (Evan Miller - How Not to Run an A/B Test). In a worst-case scenario, if you stop as soon as 5% significance appears and check after every observation, the real false-positive rate can reach 26.1% (Evan Miller - How Not to Run an A/B Test). That means roughly one in four of your "winners" is actually a loser.

The Root of the Problem: Fixed-Horizon Statistics

Most A/B testing tools default to fixed-horizon statistics, which are designed for a single analysis at a predetermined sample size. They assume you'll resist the urge to peek. But you won't, and neither will anyone else. The null hypothesis—that there's no real difference between control and variation—is the baseline, and you're trying to reject it (Spotify Confidence). In a fixed-horizon framework, every peek is a new opportunity to falsely reject that null. The conventional 5% significance level (p

A Better Way: Sequential Testing

The fix is to use a method that accounts for your peeking. Sequential testing, sometimes called always-valid inference, lets you look at your data as often as you want without inflating your error rates. The idea is simple: instead of a fixed sample size, you use a stopping rule that's designed for continuous monitoring. Evan Miller's simple sequential procedure, for example, works like this: pick a sample size N, assign users 50/50, and stop when the treatment-minus-control success count reaches 2 times the square root of N (declare a winner) or the total success count reaches N (declare no winner) (Evan Miller - Simple Sequential A/B Testing). This procedure can reduce the number of observations you need by 50% or more in some cases, and it works especially well with low conversion rates (Evan Miller - Simple Sequential A/B Testing).

Optimizely's Stats Engine is another example. It was developed with Stanford statisticians and uses sequential testing so that results are always valid no matter when you peek (Optimizely Stats Engine). When they re-ran 48,000 historical experiments, fixed-horizon statistics declared a winner or loser in 36% of tests, while the sequential engine declared winners or losers in only 22%—a 39% reduction in false conclusions (Optimizely Stats Engine). That's a massive difference in trustworthiness.

What About Bayesian? It's Not a Magic Bullet

You might have heard that Bayesian testing lets you monitor continuously without penalty. That's true up to a point. Bayesian methods output probabilities like "there is a 92% probability that Variant B is better" and allow continuous monitoring without inflating false positives (Optimizely). VWO, a popular testing platform, claims the Bayesian approach can deliver actionable results almost 50% faster than frequentist methods (VWO A/B testing guide). But here's the catch: Bayesian methods rely on prior beliefs, and if your prior is off, your probabilities are off. They also don't give you the same guarantee of controlling false positives over the long run. Sequential frequentist methods, by contrast, give you a mathematically rigorous control of the error rate. If you want a tool that's both intuitive and defensible, sequential testing is the safer bet.

How to Apply This in Your Own Tests

Here's my concrete advice. First, decide on your minimum detectable effect (MDE) and your desired power—typically 80% is the standard (Optimizely). Use a sample size calculator that accounts for your baseline conversion rate and MDE (Evan Miller - Sample Size Calculator). But don't just set a sample size and then ignore it. Instead, use a sequential testing method that lets you check results as often as you want. If your tool supports sequential testing (like Optimizely's Stats Engine), turn it on. If it doesn't, you have two options: either pre-commit to a fixed end date and resist peeking (which is hard), or use a manual sequential method like Evan Miller's.

Let's make it concrete. Suppose your baseline conversion rate is 5%, and you want to detect a 10% relative improvement (that's a 0.5 percentage point absolute increase). Using Evan Miller's rule-of-thumb, n = 16 * (sigma^2 / delta^2). For a binomial, sigma^2 = p(1-p) = 0.05 * 0.95 = 0.0475. Delta (the absolute minimum effect) is 0.005 (0.5%). So n = 16 * (0.0475 / 0.000025) = 16 * 1900 = 30,400 users per variation (Evan Miller - How Not to Run an A/B Test). That's a big number. But with sequential testing, you might be able to stop earlier if the effect is real. In one of Optimizely's re-runs, the sequential engine declared a winner or loser up to 2.5 times faster than fixed-horizon for experiments over 50,000 visitors (Optimizely Stats Engine). So you could save weeks of waiting.

Also, remember that statistical significance isn't the whole story. With huge sample sizes, even a microscopic effect becomes significant. For example, a 0.2 percentage point difference in completion rates (85.0% vs 85.2%) can be statistically significant with p = 0.03, but is that difference worth changing your design? (Nielsen Norman Group). You need to think about practical significance, not just p-values. Sequential testing helps you get to a valid answer faster, but you still need to decide whether the answer is big enough to matter.

The One Thing to Remember

Stop peeking with fixed-horizon statistics. Switch to a sequential testing method, or pre-commit to a sample size and don't look until the end. Peeking is the most common way to fool yourself into shipping a bad change. Sequential testing lets you have your cake and eat it too—you can peek all you want, and your results remain valid. Don't let another "significant" result be a false alarm.

Sources

  • Optimizely - A/B testing glossary - https://www.optimizely.com/optimization-glossary/ab-testing/
  • Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
  • Evan Miller - Sample Size Calculator - https://www.evanmiller.org/ab-testing/sample-size.html
  • Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
  • Evan Miller - Simple Sequential A/B Testing - https://www.evanmiller.org/sequential-ab-testing.html
  • Nielsen Norman Group - Practical Significance - https://www.nngroup.com/articles/practical-significance/

Share this article:

Comments (0)

No comments yet. Be the first to comment!