Skip to main content
Optimization Tools

Stop Peeking at Your A/B Test: Sequential Testing Is the Only Honest Way

Peeking at an A/B test inflates false positives up to 30%. Sequential testing, as used by Optimizely, keeps results honest. Here's why you should never stop a test early without it.

Contrarian claim: The biggest A/B testing mistake isn't bad design—it's peeking at your results before the test is done.

Most people think the hard part of A/B testing is coming up with a clever hypothesis or designing a beautiful variation. I've been there—you spend hours crafting the perfect button color, and then you can't resist checking the results every hour. But here's the thing: every time you peek and stop because you see a significant result, you're lying to yourself. The numbers you see are not the truth. Peeking at your data and stopping early can inflate your false-positive rate from a nominal 5% to as high as 30% (Optimizely). That's not a small error—that's a 6x increase in the chance that you're shipping a change that does nothing or even hurts your business. I'd argue that peeking is the single most destructive habit in optimization, and it's rampant because it feels so intuitive. We're wired to want to know the answer now, not later. But the discipline to wait is what separates professionals from amateurs.

Wait, aren't p-values the gold standard? Don't we just need p

Yes, p-values are the standard, but they're widely misunderstood. A p-value of 0.03 doesn't mean there's a 97% chance the variation is better (Optimizely). It means that if there were truly no difference, you'd see data this extreme only 3% of the time. That's subtle, and it's exactly why the misuse of p-values is so dangerous. When you peek at your data and stop as soon as p

Should I run an A/A test to make sure my setup is working?

Absolutely, and I'm surprised more people don't do this. An A/A test—where you run two identical versions—is the ultimate sanity check. If your testing tool reports a statistically significant difference between two identical pages, that's a red flag that something is broken (Optimizely). I recommend making A/A tests a regular ritual, roughly quarterly (Optimizely). It's like checking the calibration of your scale before you weigh yourself. If the scale says you gained 5 pounds when you haven't eaten anything, you don't trust the scale for your next reading. The same logic applies to your A/B testing tool. Yet I've seen teams run months of experiments on a broken setup, making decisions based on garbage data. Don't be that team.

What about sample size? How big should my test be?

This is where most people get impatient. They want to know how many visitors they need, and they want a simple number. There's a rule of thumb: n = 16 * (sigma^2 / delta^2), where delta is the minimum effect you want to detect (Evan Miller). But the real answer is: it depends on your baseline conversion rate and the minimum detectable effect (MDE) you care about (Optimizely). For a typical 80% power and 5% significance, you need a larger sample for tiny effects. The problem is that with a very large sample size, even a microscopic improvement—like a 0.03% drop in form-field errors—can produce a p-value below 0.001 (Nielsen Norman Group). That's statistically significant but practically meaningless. So don't just ask "how many visitors?" Ask "what effect size actually matters to my business?" If you can't detect a 1% lift, you might be wasting your time on a test that will never show a meaningful difference.

What is the alternative to fixed-horizon testing? Isn't there a way to look at results early without cheating?

Yes, and this is the part that gets me excited. Sequential testing is the answer. Instead of fixing a sample size in advance and then peeking, sequential testing allows you to stop early if the treatment is clearly a winner, while keeping your error rates under control. The beauty is that you can peek at any time and the results remain valid (Optimizely). Always-valid p-values and confidence intervals let you take advantage of data as fast as it becomes available (arXiv). Optimizely's Stats Engine uses this approach, and when they re-ran 48,000 historical experiments, they found that sequential testing declared winners or losers in 22% of tests, while fixed-horizon statistics declared a winner or loser in 36% (Optimizely). That means 39% fewer conclusive results—but those are the trustworthy ones. In my experience, it's better to have no answer than a false answer.

Is a multi-armed bandit better than a traditional A/B test?

This is a question I get all the time, and the answer is: it depends on your goal. A multi-armed bandit (MAB) dynamically changes the traffic split based on observed performance, so you're always sending more users to the winning variation (VWO). That's great if you want to maximize conversion during the test itself. But if you need a rigorous statistical conclusion—like for a product launch decision—a traditional A/B test with a fixed sample size and sequential testing is more appropriate. The MAB is an optimization tool, not an inference tool. It's like choosing between a race car and a family sedan: the race car is faster, but it's not great for carrying the kids. I'd use a MAB for continuous optimization, but for a one-time decision, I'd rather run a clean A/B test with sequential monitoring.

What about sample ratio mismatch (SRM)? Should I worry about that?

Yes, and this is a hidden killer. SRM is when the observed sample ratio differs from the expected ratio—like your 50/50 split becomes 48/52. It's a symptom of a data quality issue (KDD 2019). If you ignore it, you might ship a bad change because the data was skewed. I always check for SRM before I look at any other metric. If there's an SRM, I don't trust any result from that test. It's like seeing a fever—you don't ignore it; you find the infection. The same goes for novelty and primacy effects. Novelty is the initial boost from something new, and primacy is the growing engagement as users learn (arXiv). These can make a variation look good or bad in the short term, but the effect fades. So don't be fooled by a spike in conversions in the first week.

Bottom line

If you take one thing from this, let it be this: stop peeking and start using sequential testing. It's the only honest way to run an A/B test in the modern era. Pre-commit to a fixed end date or use a tool that supports sequential analysis. And for the love of everything, run an A/A test every quarter to make sure your setup is sound. Your future self—and your revenue—will thank you.

Sources

  • Optimizely A/B testing glossary - https://www.optimizely.com/optimization-glossary/ab-testing/
  • Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
  • Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
  • Optimizely A/A testing glossary - https://www.optimizely.com/optimization-glossary/aa-testing/
  • Nielsen Norman Group - Practical Significance - https://www.nngroup.com/articles/practical-significance/
  • arXiv - Always Valid Inference (Johari et al.) - https://arxiv.org/abs/1512.04922

Share this article:

Comments (0)

No comments yet. Be the first to comment!