Skip to main content
Test Design

A/B Testing: The Art of Knowing When to Stop (and When to Keep Going)

I've seen too many teams call a test early and ship a dud. Here's how to design an A/B test that actually gives you a clear answer—without the statistical hand-waving.

How do I know when my A/B test has run long enough?

You've launched your test. The dashboard shows a 12% lift after three days. Should you call it? I've been there. And I've made the mistake of calling it too early. The lift vanished. The feature shipped anyway. Nobody was happy.

The problem isn't your dashboard. It's your test design. Before you start, you need a plan. Not a vague idea. A specific, written-down plan. In this post, I'll show you how to build one.

Picture this: you're a PM at a mid-sized e-commerce site

You want to test a new checkout button color. Your baseline conversion rate is 3%. You think the new button will lift conversion by at least 10% relative—so from 3% to 3.3%. That's your minimum detectable effect. You get 10,000 visitors a day. You're eager. What do you do?

Wrong answer: launch and check daily. Right answer: design the test first. Let's design it.

First, state your hypotheses and pick your significance level

Every test starts with a null hypothesis: the new button has no effect. The alternative: it does. You'll use a significance level (alpha) of 0.05. That's the conventional threshold. It means you accept a 5% chance of a false positive—a Type I error. You also want 80% power. Power is the probability of detecting a real effect if it exists. With these numbers, you can calculate the sample size you need.

Next, calculate the sample size—and stick to it

Using the baseline rate (3%) and your minimum detectable effect (10% relative), a sample size calculator tells you that you need about 47,000 visitors per variation. That's 94,000 total. At 10,000 visitors per day, you need 9.4 days. Round up to 10 days. This is your pre-committed end date. No stopping early, no matter how tempting.

Why? Because peeking inflates false positives. If you check results daily and stop when you see significance, your false-positive rate can jump from 5% to as high as 30%. In a worst-case scenario with continuous monitoring, it can reach 26.1%. That means one in four 'wins' is actually noise. Not good.

Here's a concrete example: I once ran a test where the early lift was 15% after two days. We had planned for 14 days. By day 14, the lift was 1.2%—not significant. If we had stopped early, we would have shipped a feature that did nothing.

But what if you want to stop early? Consider sequential testing

You might think: 'But I want to see results early.' I get it. That's why sequential testing exists. Methods like Evan Miller's simple sequential procedure let you stop early if the treatment is clearly winning, while controlling false positives. For example, you choose a sample size N and stop when the treatment-minus-control success count reaches 2 times the square root of N. In some cases, this can reduce required observations by 50% or more. Optimizely's Stats Engine uses always-valid p-values, so you can peek anytime without inflating errors. But even with sequential methods, you need to decide on a framework before you start. Don't mix and match.

Check for sample ratio mismatch and other data quality issues

Before you trust any result, verify your traffic split. If you expect a 50/50 split but see 52/48, you have a sample ratio mismatch (SRM). An SRM is like a fever—it signals a data quality problem. Ignoring it can make a bad change look good. Run an A/A test quarterly to validate your setup. If your A/A test shows a significant difference, your implementation is broken.

Interpret the results: statistical vs. practical significance

Suppose after 10 days you see a 0.2 percentage point lift (3.0% to 3.2%) with p = 0.03. Statistically significant, right? But is it practically significant? A 0.2 point lift on 3% is a 6.7% relative improvement. On your revenue, that might be meaningful. But consider the Nielsen Norman Group example: a 0.2 point difference in completion rate (85.0% vs 85.2%) was statistically significant (p = 0.03) yet too small to justify a design change. So ask: does this lift justify the engineering cost? Only you can decide, but don't let p

Beware of novelty effects and interactions

Even with a perfect design, users might react to novelty. Novelty effects can inflate early results and then fade. If your test runs only a few days, you might mistake novelty for a real lift. Running for at least one full business cycle (typically two weeks) helps. Also, if you're running multiple experiments simultaneously, watch for interactions. At Bing, with over 250 concurrent experiments, statistical interactions are common (Kohavi et al.). Your mid-sized site might not have that scale, but if you're running a site-wide banner test and a checkout test at the same time, they could interfere.

My recommendation: pre-commit, then don't touch it

Design your test with a fixed sample size, a pre-committed end date, and a clear decision rule. Use sequential testing only if you need early stopping and you've chosen a valid method. Run A/A tests quarterly. Check for SRM. And when you get a significant result, ask whether it's practically significant. This isn't glamorous, but it's how you avoid shipping noise. As Jim Barksdale said, 'If we have data, let's look at data. If all we have are opinions, let's go with mine.' In A/B testing, the data only speaks clearly if you design the test to listen.

Sources

  • Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
  • Evan Miller - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
  • KDD 2019 - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom
  • Nielsen Norman Group - https://www.nngroup.com/articles/practical-significance/
  • Kohavi et al. - https://exp-platform.com/large-scale/

Share this article:

Comments (0)

No comments yet. Be the first to comment!