How do I design an A/B test that gives me a trustworthy answer? That's the question you typed into Google, and I'm going to answer it with a concrete scenario, not a lecture. Imagine you are the growth lead at a mid-sized e-commerce site. Your checkout page converts at 3.0%. Your designer has a new layout they swear will lift conversions. Your boss wants results by Friday. You have about 20,000 visitors per day hitting that page. What do you do?
Start with a hypothesis, not a hunch
You need to state two competing claims before you touch a single line of code. The null hypothesis says there's no real difference between your control (the current checkout) and your variation (the new layout) — any gap you see is just noise (Spotify Confidence). The alternative hypothesis says there is a difference. This isn't academic ceremony. If you skip this step, you'll end up chasing whatever number looks shiniest on day three.
Pick your significance level and power before you launch
Set alpha at 0.05. That means when there's genuinely no difference, you'll falsely declare a winner 5% of the time (NIST e-Handbook). Set power at 0.80 — an 80% chance of detecting a real effect if one exists. These aren't arbitrary; they're the standard defaults (Optimizely). Now decide your minimum detectable effect (MDE). If your baseline is 3.0% and you only care about changes of 0.3 percentage points or more, say so. The MDE is the smallest effect your test can reliably catch; anything smaller is indistinguishable from noise (Evan Miller - Sample Size Calculator).
Calculate your sample size and don't budge
Here's where most teams go wrong. You need to figure out how many visitors per variation you need. The formula depends on your baseline rate, your MDE, your alpha, and your power. A useful rule of thumb is n = 16 × (σ² / δ²), where δ is your minimum effect and σ² is the variance — for a conversion rate, that's p(1-p) (Evan Miller - How Not to Run an A/B Test). With a 3% baseline and a 0.3 percentage point MDE, you're looking at roughly 16 × (0.03 × 0.97) / (0.003²) ≈ 51,733 visitors per variation. That's over 100,000 total. At 20,000 visitors per day, you need about five days. Write that number down. Tape it to your monitor.
The peeking trap that will ruin your test
Now the hard part: not looking. If you check your results daily and stop as soon as you see p
Quick tip: Pre-commit to a fixed end date. If you absolutely must monitor, use a sequential testing method that adjusts its thresholds for multiple looks (Optimizely).
Check for sample ratio mismatch before you trust anything
You planned a 50/50 split. If your analytics show 51/49, you have a sample ratio mismatch (SRM). Think of an SRM as a fever: it's a symptom of a data quality problem, not the disease itself (KDD 2019 - Diagnosing Sample Ratio Mismatch). Ignore it and you might ship a bad change thinking it's good, or kill a good one thinking it's bad. Run an A/A test quarterly to validate your setup. If your A/A test shows a significant difference between two identical versions, your implementation is broken (Optimizely A/A testing glossary).
Choose your statistical framework wisely
You have three real options. Here's how they compare for your checkout test:
| Framework | How it handles peeking | Sample size efficiency | Best for |
|---|---|---|---|
| Fixed-horizon frequentist | Peeking inflates false positives; must wait until pre-set N | Baseline; no reduction | Simple, high-traffic tests where you can wait |
| Sequential testing (e.g., SPRT, always-valid p-values) | Allows continuous monitoring with valid inference | Can reduce observations by 50% or more in some cases | Tests where early stopping saves real money |
| Bayesian | Continuous monitoring without inflating false positives | VWO reports actionable results almost 50% faster | Stakeholders who want probability statements, not p-values |
My recommendation? For your checkout test, use a sequential method. Evan Miller's simple sequential procedure lets you stop early if the treatment-minus-control success count reaches 2 times the square root of N, or declare no winner if total successes hit N (Evan Miller - Simple Sequential A/B Testing). It works extremely well with low conversion rates — exactly your situation. And it can cut the observations you need by 50% or more. That means instead of five days, you might get your answer in two or three. If your organization is allergic to p-values, Bayesian gives you a direct probability statement like "there's a 92% chance the new layout is better" (Optimizely). Either way, stop pretending you can just eyeball the dashboard.
Don't confuse statistical significance with practical significance
Suppose your test finishes and the new layout lifts conversion from 3.00% to 3.02%. With enough traffic, that difference could be statistically significant. So what? A 0.2 percentage point difference (85.0% versus 85.2%) can be statistically significant at p = 0.03 yet too small to justify a redesign (Nielsen Norman Group). Ask yourself: does this effect size matter to the business? At Bing, a 1% revenue improvement equals about $10 million annually in the US (Kohavi et al.). Your 0.02 percentage point lift on a $5 million checkout flow is $1,000. Not worth the engineering risk.
Finally, remember that less than a third of ideas tested at Bing move the metrics they were designed to improve (Bing). Your designer's layout is probably not a winner. That's fine. The goal isn't to validate your intuition; it's to find the truth cheaply and move on.
The one thing to remember
Design your test like you're going to lose. Pre-commit your sample size, your MDE, and your end date. Then don't peek. If you do peek, use a sequential method that's built for it. Everything else — the calculator, the dashboard, the fancy Bayesian posterior — is just machinery. The discipline is deciding what would change your mind before you see the data, and then actually waiting for it.
Sources
- Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
- Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Evan Miller - Simple Sequential A/B Testing - https://www.evanmiller.org/sequential-ab-testing.html
- KDD 2019 - Diagnosing Sample Ratio Mismatch - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom
- Nielsen Norman Group - Practical Significance - https://www.nngroup.com/articles/practical-significance/
- Spotify Confidence - https://confidence.spotify.com/glossary/frequentist-ab-testing
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!