You check your A/B test on day three, and the p-value is 0.04. Tempting, right? But here's the cold truth: if you peek at your results daily and stop when significance appears, your false-positive rate can balloon from 5% to as high as 30% (Optimizely). That's not a rounding error—it's a process failure. You're not testing; you're gambling.
The Peeking Problem
Why is peeking so dangerous? Because every time you look and consider stopping, you're running a separate hypothesis test on the same data. The math doesn't forgive. Evan Miller, who's written extensively on this, shows that if you peek ten times at an experiment, what you think is 1% significance is actually only 5% (Evan Miller). And if you stop at the first sign of 5% significance after every single observation, the real false-positive rate reaches 26.1% (Evan Miller). That means over a quarter of your 'winners' are pure noise.
You might think, 'I'll just set a fixed sample size in advance and not look until then.' That's the classic frequentist approach, and it works—if you have the discipline. But here's the catch: in fast-moving product teams, discipline is rare. Business stakeholders want results now. And peeking isn't just about ego; it's about speed. Waiting for a pre-calculated sample size can feel like an eternity when your CEO is asking for a decision.
Sequential Testing: The Safer Way to Speed
What if you could look at your data as often as you like, without inflating your error rates? That's the promise of sequential testing. Instead of a fixed sample size, sequential methods adjust your thresholds as you accumulate data, allowing you to stop early if there's a clear winner—while keeping your false-positive rate under control. Evan Miller's simple sequential procedure is a great starting point: pick a sample size N, split traffic 50/50, and stop when the treatment-minus-control success count reaches 2 times the square root of N (declaring a winner), or when total successes hit N (declaring no winner) (Evan Miller). It's elegant, and in many cases it can cut the number of observations needed by 50% or more, especially with low conversion rates (Evan Miller).
Optimizely has built its Stats Engine around this idea. Instead of forcing you to pre-commit to an MDE, their sequential method gives you 'always-valid' p-values—you can peek any time and trust the result (Optimizely Stats Engine). In a re-analysis of 48,000 historical experiments, fixed-horizon statistics declared a winner or loser in 36% of tests, while the sequential engine only called it in 22% (Optimizely Stats Engine). That might sound like fewer wins, but it's actually fewer false positives. The sequential engine was uncovering 39% fewer 'conclusive' results—results that were likely statistical mirages.
Why Fixed-Horizon Tests Are a Trap
The problem with fixed-horizon tests isn't just peeking; it's the false sense of precision. You calculate a sample size based on a minimum detectable effect (MDE) you guessed at, then you run the test for that exact duration. But business conditions change, traffic fluctuates, and your original MDE might be unrealistic. If you set your MDE too large, you'll miss real improvements; too small, and you'll waste weeks on noise.
Even if you follow the rules, fixed-horizon tests have another flaw: they treat every day as equal. But in reality, your conversion rate on Mondays might be different from Saturdays. Sequential methods are more robust to these variations because they don't assume a fixed end point.
And let's talk about practical significance. A large sample size can make tiny effects look statistically significant. For example, a 0.2 percentage point difference in completion rate (85.0% vs. 85.2%) can be statistically significant with p=0.03, but it's probably not worth redesigning your checkout flow (Nielsen Norman Group). Sequential testing helps you focus on effects that matter, because you can stop as soon as you have enough evidence—not after you've collected a mountain of data that makes everything significant.
Putting Sequential Testing into Practice
You might be thinking, 'Okay, but how do I actually do this?' Most modern A/B testing platforms—like Optimizely—have sequential methods built in. If you're using a platform that doesn't, you can implement Evan Miller's procedure yourself. Or, if you're a Bayesian, you can use continuous monitoring without inflating false positives (Optimizely). The key is to stop treating your A/B test like a fixed-duration bake-off and start treating it like a decision process that can end as soon as you have enough evidence.
Let me give you a concrete example. Suppose you're testing a new headline on your landing page. Your current conversion rate is 5%. You want to detect a 0.5 percentage point improvement (a relative lift of 10%). With a fixed-horizon test, you'd calculate a sample size of about 16 * (p*(1-p)/delta^2), where delta is 0.005. That's 16 * (0.05*0.95)/(0.005^2) = 16 * 0.0475/0.000025 = 30,400 visitors per variation, or 60,800 total (Evan Miller). If you're getting 1,000 visitors a day, that's over two months. But with sequential testing, you might stop in half the time if the new headline is clearly better—or clearly not.
What I'd Actually Do
Here's my blunt recommendation: ditch fixed-horizon testing for anything that matters. Switch to a platform that supports sequential testing, or implement a simple sequential procedure yourself. Yes, it's a change in mindset, but the payoff is real: you'll stop wasting time on doomed tests, and you'll catch true winners faster.
Start by auditing your current testing process. Are you peeking? If so, stop. Then, for your next test, use a sequential method. If you're on Optimizely, turn on Stats Engine. If you're on a platform that doesn't have it, use Evan Miller's procedure as a fallback. And always run an A/A test first to make sure your setup is sound (Optimizely A/A testing glossary).
The days of 'set it and forget it' are over. Sequential testing is not just a statistical nicety; it's a competitive advantage. You'll make better decisions, faster, and you'll sleep better knowing your 'winners' are real.
Sources
- Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
- Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Evan Miller - Simple Sequential A/B Testing - https://www.evanmiller.org/sequential-ab-testing.html
- Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
- Nielsen Norman Group - Practical Significance - https://www.nngroup.com/articles/practical-significance/
- Optimizely A/A testing glossary - https://www.optimizely.com/optimization-glossary/aa-testing/
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!