Skip to main content
Statistical Analysis

Stop Peeking: Why Your A/B Test Results Are a Statistical Illusion

Peeking at A/B test results can inflate false positives to 30%. Precommit to a sample size or use sequential testing to make valid decisions.

Who This Is For

If you’re a product manager, marketer, or data scientist who has ever felt that familiar itch to check your A/B test results before the scheduled end date, this article is for you. We’ve all been there: the dashboard shows a promising lift after just a few days, and the temptation to call it early is almost irresistible. But here’s the contrarian truth: peeking at your experiment and stopping as soon as significance appears can inflate your false-positive rate from a nominal 5% to as high as 30% (Optimizely). That means nearly one in three of your “winners” could be pure noise. In this walkthrough, we’ll show you how to run A/B tests that you can actually trust, by pre-committing to a sample size or using sequential methods that let you peek without penalty.

1. Start with a Clear Hypothesis and a Pre-Commitment

Before you launch any test, you need a null hypothesis—the assumption that there is no real difference between your control and variation (Spotify Confidence). Then, decide on a significance level (usually 0.05) and the statistical power you want, typically 80% (Optimizely). The power is your chance of detecting a real effect if one exists, and it’s tied to your minimum detectable effect (MDE)—the smallest lift you care about. The key move is to decide on your sample size in advance, based on your baseline conversion rate and your MDE (Evan Miller). This pre-commitment is your defense against the urge to peek. Without it, you’re vulnerable to the classic mistake: checking results daily and stopping as soon as significance appears, which can push your false-positive rate to 30% (Optimizely).

2. Calculate Sample Size Using a Rule of Thumb

You don’t need a PhD to estimate how many visitors you need. A rule-of-thumb formula is n = 16 * (sigma-squared / delta-squared), where delta is the minimum effect you wish to detect and sigma-squared is the expected variance of your metric (Evan Miller). For a conversion rate, the variance is p(1-p), where p is your baseline conversion rate. Let’s make it concrete: suppose your current checkout conversion is 10% (p = 0.10), so the variance is 0.09. If you want to detect a 1% absolute lift (delta = 0.01), the formula gives n = 16 * (0.09 / 0.0001) = 14,400 visitors per variation. That’s a big number, but it’s the cost of not fooling yourself. If you have 1,000 visitors per day, you’d need about 29 days to reach that sample size (Optimizely). Many teams underestimate how long tests need to run, and then they peek because they’re impatient.

3. Choose a Fixed Horizon or Use Sequential Testing

Once you’ve set your sample size, you have two valid paths. The classic approach is to run the test for a fixed number of visitors or a fixed duration, and not look at the results until the end. This is the fixed-horizon method that most traditional statistics assume. The alternative is sequential testing, which allows you to look at the data as it comes in and stop early if the treatment is clearly a winner—all without inflating your false-positive rate (Evan Miller). Sequential methods like the Sequential Probability Ratio Test (SPRT) can potentially reduce required sample sizes and save resources (Optimizely). In fact, when Optimizely re-ran 48,000 historical experiments, their sequential Stats Engine declared winners or losers in only 22% of tests, compared to 36% for fixed-horizon statistics—revealing 39% fewer conclusive results, meaning many apparent wins were likely false positives (Optimizely).

4. Understand the Numbers: p-values, Confidence Intervals, and Power

You’ve likely heard of p-values and confidence intervals, but do you really know what they mean? A p-value is the probability of observing data as extreme as what you saw, assuming the treatment has no effect. A p-value of 0.03 means such a result would occur by chance 3% of the time if there were no real difference (Optimizely). It does NOT mean there’s a 97% chance the treatment works. A 95% confidence interval, on the other hand, is a range that, if you repeated the experiment many times, would contain the true effect 95% of the time (Optimizely). Power is your chance of correctly rejecting a false null hypothesis—the probability of not missing a real effect (Stat Trek). If your test is underpowered, you may fail to detect a genuine improvement, which is a Type II error (Spotify Confidence). That’s why you set power to 80% before you start.

5. The Peeking Problem: Why It’s So Dangerous

Peeking is the habit of checking your results multiple times during the experiment and stopping as soon as you see significance. The math is brutal: if you peek at an ongoing experiment ten times, what you think is 1% significance is actually just 5% significance (Evan Miller). In a worst-case scenario where you stop as soon as 5% significance appears, with a significance test run after every observation, the actual false-positive rate can reach 26.1% (Evan Miller). That’s more than one in four of your “wins” being false. The defense is simple: pre-commit to a fixed end date or use sequential testing that adjusts thresholds for multiple looks (Optimizely). Peeking at the data is acceptable as long as you restrain yourself from stopping the experiment before it has run its course (Evan Miller). So go ahead and look, but don’t act on what you see until the pre-determined end.

6. Watch Out: Sample Ratio Mismatch and Other Data Quality Issues

Even if you follow all the statistical rules, your test can be invalidated by data quality issues. A Sample Ratio Mismatch (SRM) occurs when the observed sample ratio in an experiment differs from the expected ratio—like a fever that signals an underlying illness (KDD 2019). Ignoring an SRM without knowing its root cause may result in a bad product modification appearing to be good and being shipped to users, or vice versa (KDD 2019). So before you celebrate a win, check that your traffic split matches what you set. Also, be wary of novelty and primacy effects: novelty describes the initial excitement users feel about a new feature, which can fade over time, while primacy describes growing engagement as users adopt the innovation (arXiv). These can distort early results, which is another reason not to peek too soon.

7. Consider Alternatives: A/A Tests, CUPED, and Bayesian Methods

If you want to validate your testing setup, run an A/A test—comparing two identical versions—to ensure your system isn’t producing false positives. Optimizely recommends doing this quarterly (Optimizely). If your A/A test reports a statistically significant difference between identical versions, that’s a sign your implementation needs checking (Optimizely). To increase power without a longer test, you can use CUPED (Controlled-experiment Using Pre-Experiment Data), which adjusts your metrics using pre-experiment data to reduce variability. At Bing, three experiments using CUPED showed variance reductions of 45%, 52%, and 49% (Deng et al.). That effectively halves the required sample size or duration. Finally, if you’re tired of p-values that confuse everyone, Bayesian testing might appeal to you—it outputs probabilities like “there is a 92% probability that Variant B is better” and allows continuous monitoring without inflating false positives (Optimizely). But Bayesian methods come with their own assumptions, so choose the framework that fits your team’s decision culture.

What Can Go Wrong: The Seduction of Tiny Effects

One more trap: statistical significance does not equal practical significance. With a very large sample size, even a microscopic improvement—like a 0.03% drop in form-field errors—can produce a p-value below 0.001 (Nielsen Norman Group). That doesn’t mean you should redesign your form for such a tiny gain. Conversely, an effect that is not statistically significant may still warrant action if it’s large and meaningful. For example, an 80% drop in task completion observed in a small-sample usability test is alarming even if it doesn’t reach p

Bottom Line

The single best move you can make is to pre-commit to a sample size before you launch, based on your baseline conversion rate and minimum detectable effect, and then resist the urge to stop early—or use a sequential testing method that lets you peek without penalty. Your future self will thank you for avoiding the 30% false-positive trap.

Sources

  • Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
  • Evan Miller - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
  • Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
  • Nielsen Norman Group - https://www.nngroup.com/articles/practical-significance/
  • Deng et al. - https://exp-platform.com/Documents/2013-02-CUPED-ImprovingSensitivityOfControlledExperiments.pdf
  • KDD 2019 - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom

Share this article:

Comments (0)

No comments yet. Be the first to comment!