Skip to main content
Optimization Tools

Stop Peeking at Your A/B Test: Why Waiting Is the Best Optimization Tool

Peeking at A/B test results destroys their validity. Learn the right way to run tests, when to stop, and how to get reliable answers faster.

You think checking your A/B test every morning and stopping as soon as it hits 95% significance is being agile. You're wrong. That habit is quietly turning your experiments into coin flips. I've seen it happen too many times: a marketer declares victory on Monday, ships the so-called winner, and then wonders why conversion rates mysteriously drop the following week. The culprit isn't a bad idea—it's bad statistics. And the fix isn't a fancier tool; it's understanding what your optimization tool is actually telling you.

This article is for anyone running A/B tests on landing pages, product features, or marketing campaigns. If you've ever felt that nervous itch to check results before the test ends, this is for you. I'm going to walk you through a practical, step-by-step approach to running tests that you can actually trust, without needing a PhD in statistics.

1. Know What You're Testing For: The Null Hypothesis Isn't Your Enemy

Before you write a line of code, get clear on the null hypothesis. It's simply the assumption that there's no real difference between your control and variation—any difference you see is just random noise (Spotify Confidence). I used to gloss over this, thinking it was academic fluff. But it's the foundation for everything else. When you set up an A/B test, you're not trying to prove your variation is better; you're trying to gather enough evidence to reject the null. That shift in mindset changes how you interpret every metric.

Here's a concrete example: you're testing a new headline. Your control converts at 5%, your variation at 5.2%. That 0.2% lift might feel good, but the null hypothesis says it's likely just chance. Your job is to determine whether the evidence is strong enough to reject that assumption.

2. Decide Your Sample Size Before You Start: Don't Guess

This is the single most important step, and most people skip it. You need to calculate how many visitors your test requires before you let it run. The math is straightforward: your baseline conversion rate, the minimum effect you care about, your desired significance level (usually 0.05), and statistical power (commonly 80%) (Optimizely sample size guide). The rule of thumb I use: n = 16 times (sigma-squared / delta-squared), where delta is the smallest effect you want to detect and sigma-squared is variance (Evan Miller - How Not to Run an A/B Test). For conversion rates, variance is p(1-p).

For example, if your baseline conversion is 10% and you want to detect a 1% absolute lift, delta is 0.01, and sigma-squared is 0.1 * 0.9 = 0.09. Plug it in: 16 * (0.09 / 0.0001) = 14,400 visitors per variation. That's a lot more than most people expect. But without that number, you're flying blind.

3. Pre-Commit to a Duration: The Antidote to Peeking

Once you know your sample size, set a fixed end date. Then don't touch the test until that date. Peeking—checking results daily and stopping when significance appears—inflates your false-positive rate from 5% to as high as 30% (Optimizely (A/B testing)). That's not a typo. If you peek ten times, what you think is 1% significance is actually 5% (Evan Miller - How Not to Run an A/B Test). In a worst-case scenario, stopping as soon as 5% appears with a test after every observation gives a false-positive rate of 26.1% (Evan Miller - How Not to Run an A/B Test).

I know it's hard. I've been there, refreshing the dashboard like it's a stock ticker. But discipline here is what separates real optimization from superstition.

4. Use Sequential Testing: The Smart Way to Check Early

If you absolutely can't wait, use a tool that supports sequential testing. Sequential methods let you peek anytime without ruining your results (Evan Miller - Simple Sequential A/B Testing). Optimizely's Stats Engine, for example, uses sequential testing and can call a winner or loser up to 2.5 times faster than fixed-horizon statistics for experiments over 50,000 visitors (Optimizely Stats Engine). Netflix uses sequential canary tests to detect regressions in streaming metrics like PlayDelay (arXiv - Netflix canary testing (Lindon et al.)).

Here's a simple sequential procedure I like: pick a sample size N, assign visitors 50/50, and stop when the treatment-minus-control success count reaches 2 times the square root of N (winner) or total successes reach N (no winner) (Evan Miller - Simple Sequential A/B Testing). This can cut required observations by 50% or more in some cases (Evan Miller - Simple Sequential A/B Testing). That's a game-changer for busy teams.

5. Watch Out for Sample Ratio Mismatch (SRM)

Even with perfect math, your data can lie. A Sample Ratio Mismatch (SRM) is when the observed ratio of visitors in control vs. variation differs from what you set. It's a red flag that something's wrong with your implementation (KDD 2019 - Diagnosing Sample Ratio Mismatch). Ignoring it can make a bad feature look good, or vice versa (KDD 2019 - Diagnosing Sample Ratio Mismatch).

I once had a test where the variation got 60% of traffic due to a caching bug. The results looked fantastic—until I noticed the mismatch. Always check for SRM before trusting any result. It's the first thing I look at after a test ends.

6. Don't Be Fooled by Statistical Significance Alone

Statistical significance (p < 0.05) only tells you that the result is unlikely to be due to chance. It says nothing about whether the effect is big enough to matter (Nielsen Norman Group - Practical Significance). A completion-rate difference of 0.2 percentage points (85.0% vs. 85.2%) can be statistically significant with p = 0.03, but it's probably not worth changing your design (Nielsen Norman Group - Practical Significance).

On the flip side, a non-significant result might still be actionable. If a usability test shows an 80% drop in task completion, that's practically significant even if the p-value is high (Nielsen Norman Group - Practical Significance). Use your head, not just the p-value.

7. Consider CUPED: Make Your Tests More Sensitive

If you're impatient, CUPED (Controlled-experiment Using Pre-Experiment Data) can help. It uses pre-experiment data to reduce metric variability, giving you the same power with fewer users or less time (Deng et al. - CUPED (WSDM 2013)). At Bing, CUPED reduced variance by 45% to 52%, effectively halving the required sample size (Deng et al. - CUPED (WSDM 2013)). That's a huge win for speed.

The best covariate is usually the same metric measured before the experiment (Deng et al. - CUPED (WSDM 2013)). Check if your tool supports it—it's a free speed boost.

What Can Go Wrong: The Peeking Trap

I can't stress this enough: if you peek and stop early, you're not just wasting time—you're making decisions on garbage data. The false-positive rate can balloon to 26.1% (Evan Miller - How Not to Run an A/B Test). I've seen teams ship features that actually hurt revenue because they couldn't wait two more weeks. Don't be that team.

Quick tip: Before you launch any test, write down the sample size, duration, and stopping rule on a sticky note. When temptation hits, look at that note.

Sources

  • Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
  • Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
  • Optimizely sample size guide - https://www.optimizely.com/insights/blog/how-to-calculate-sample-size-of-ab-tests/
  • Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
  • Deng et al. - CUPED - https://exp-platform.com/Documents/2013-02-CUPED-ImprovingSensitivityOfControlledExperiments.pdf
  • KDD 2019 - Diagnosing Sample Ratio Mismatch - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom

Share this article:

Comments (0)

No comments yet. Be the first to comment!