Skip to content
Can Elmas

CRO · 8 min read

A/B Test Sample Size and Significance: When Is a Test Really Done?

TL;DR

A test is done when it reaches the sample size you calculated before launch, over whole weeks, not when the tool shows a winner. Set the minimum detectable effect first, since it drives sample size, then judge results by the confidence interval. Stopping at the first significant reading turns coin flips into false wins.

· Published · Updated

A test is done when it reaches the sample size you calculated before launch, over whole weekly cycles, and not a day earlier because the dashboard turned green. Most tests stopped early are coin flips reported as wins. Set the minimum detectable effect first, let it fix the sample size and duration, and read the result from the confidence interval rather than the winner badge.

Why peeking at results creates false winners

A standard A/B test, the fixed-sample kind most calculators assume, accepts a 5% false positive rate on one condition: you evaluate the result once, at the planned sample size. Every extra look is another chance for random noise to cross the significance line.

Early in a test, conversion rates swing hard. With a few hundred conversions per variant, a couple of lucky days can open a gap that looks decisive. If you check daily and stop the first time the tool says “significant,” you’re no longer running one test at 5% risk. You’re taking dozens of draws, and the real false positive rate climbs far above 5%. Run an A/A test, two identical variants, with daily peeking and there’s a real chance you’ll crown a “winner” at some point, even though nothing differs.

Early stopping does a second kind of damage: the wins it produces are inflated. A test stopped the moment noise pushed it over the line captures the lift at its peak. Ship it, and the real effect is usually smaller, sometimes zero. That’s how teams stack “winning” tests and never see the combined lift in revenue.

Two rules fix most of it:

  1. Calculate the sample size before launch and make no decision before you reach it.
  2. Monitor for breakage, not for winners. Mid-test checks look for tracking errors, a broken variant or an uneven traffic split.

Some testing tools use sequential statistics built for continuous monitoring. If yours does, follow its stopping rules exactly; if unsure, assume fixed-sample.

Minimum detectable effect and why it drives everything

The minimum detectable effect (MDE) is the smallest true lift your test is designed to reliably detect. It looks like a statistics input, but it’s a business decision: what’s the smallest improvement worth shipping and worth the traffic it costs to prove?

It drives everything because sample size scales with the inverse square of the effect. Halve the MDE and you need roughly four times the visitors. A test built to detect a 5% relative lift needs about 15 times the traffic of one built to detect 20%.

Settle three things before you pick a number:

  • Relative or absolute. A 10% relative lift on a 3% conversion rate means 3.0% to 3.3%, an absolute change of 0.3 percentage points. Calculators ask for one or the other, and mixing them up produces a sample size that’s badly wrong.
  • One primary metric. Conversion to the next step needs far less traffic than revenue per visitor, because order values vary so much. Keep revenue, average order value and refunds as guardrails you check, not the metric that decides.
  • What the change can plausibly move. A button color won’t produce a 20% lift. A new offer, a rebuilt pricing page or a shorter signup form might. Small changes need huge samples; bold changes can be tested on modest traffic.

If the MDE your traffic supports is bigger than any realistic change could produce, don’t run the test. CRO for low-traffic sites covers what to do instead.

Calculating sample size and test duration

You need four inputs: baseline conversion rate, MDE, significance level (usually 5%, meaning 95% confidence) and statistical power (usually 80%, meaning an 80% chance of detecting a real effect the size of your MDE). Any reputable sample size calculator takes these. For a quick sanity check on conversion rates, a common approximation is:

Visitors per variant ≈ 16 × p × (1 − p) ÷ d²

Here p is the baseline conversion rate and d is the absolute difference you want to detect, at 95% confidence and 80% power. With a 3% baseline and a 10% relative MDE (d = 0.003), that’s 16 × 0.03 × 0.97 ÷ 0.000009, or about 52,000 per variant. An exact calculator gives about 53,000.

Duration is total sample divided by daily eligible visitors, rounded up to whole weeks. Eligible means visitors who actually reach the tested page or step, not sitewide sessions.

A hypothetical page with 5,000 eligible visitors a day, a 3% baseline and two variants:

Relative MDEConversion rate changeVisitors per variantDays of traffic neededPlanned duration
5%3.00% → 3.15%~208,000~8412 weeks: too long, test a bolder change
10%3.00% → 3.30%~53,000~224 weeks
20%3.00% → 3.60%~14,000~62 weeks (floor)
30%3.00% → 3.90%~6,500~32 weeks (floor)

Note the 10% row: 22 days rounds up to four full weeks, not three. Each extra variant adds another full sample, so stick to A/B unless traffic is plentiful.

Set a ceiling of six to eight weeks. The longer a test runs, the more cookie churn and seasonal drift contaminate it; some browsers cap client-side cookie lifetimes, so returning visitors can switch variants mid-test.

Statistical significance and confidence intervals in plain language

Significance answers a narrow question: if the variant truly made no difference, how surprising is this result? A p-value below 0.05 means a gap at least this large would show up less than 5% of the time by chance alone. It does not mean there’s a 95% chance the variant is better, and it tells you nothing about how big the lift is.

The confidence interval is more useful. It’s the range of true effects consistent with your data. Two hypothetical results, each from the 53,000-per-variant test above:

Test ATest B
Observed relative lift+12%+5%
95% confidence intervalabout +5% to +19%about −2% to +12%
Significant at 95%?YesNo
How to read itReal lift, likely smaller than +12%Inconclusive: small win, nothing or small loss

Four ways an interval can land:

  • Excludes zero, and the low end is worth having: ship it.
  • Excludes zero, but the low end is trivial: ship if it’s cheap to maintain, and don’t forecast from the point estimate.
  • Includes zero, narrow around it: the change barely matters either way; keep the simpler version.
  • Includes zero, wide: the test was underpowered and you learned little.

When I forecast revenue from a win, I use the lower bound, not the headline lift. And remember what 80% power implies: one in five real improvements of MDE size will come back inconclusive. “Not significant” is not proof of no effect.

Your tool’s winner badge may rest on a sound statistic, but it doesn’t know your planned sample or business cycle. Check the interval against the plan before calling anything.

Novelty effects, seasonality and full business cycles

Sample size tells you how many visitors you need. It doesn’t guarantee they’re representative.

Day-of-week patterns. B2B traffic on a Tuesday behaves differently from Saturday traffic, and ecommerce shoppers around paydays differ from mid-month browsers. Run whole weeks, start and stop on the same weekday, and hold a two-week floor even if the sample fills sooner.

Novelty and change aversion. Returning visitors react to anything new. A redesigned navigation can get extra clicks for a week because it’s different, or fewer because people can’t find what they’re used to. Compare the lift week by week: a big first week followed by a flat second one usually means novelty. For changes regulars will notice, new-visitor results are the cleaner read.

Outside events. A sale, a big email send or a shift in paid campaigns changes who arrives. Don’t launch into a promotion, and log every marketing event during the test.

Sample ratio mismatch. On a 50/50 split, a 51/49 result across 100,000 visitors is far outside what chance produces, and a chi-square test on the counts will flag it. The usual causes are redirect bugs, bot filtering or uneven tracking. Until you find the cause, the results can’t be trusted.

Segment results: useful or misleading

After a flat test, it’s tempting to slice by device, new versus returning, traffic source or country. Somewhere a segment will show a winner. That’s arithmetic, not insight: check 20 segments at 95% confidence and you should expect about one false “significant” result even when the variant does nothing. Segments also hold a fraction of the sample, so their intervals are much wider than the overall one.

  • Pre-register one or two segments with a reason, such as “the change only touches the mobile menu, so the effect should show on mobile.”
  • Treat unplanned segment wins as hypotheses for the next test, never as results to ship.
  • If the change is for one segment, target the test to it from the start and size the sample on that segment’s traffic.

Feeding those hypotheses back into a ranked backlog is covered in the growth experimentation process.

A pre-test checklist

  • Hypothesis written: the change, the metric it should move and why
  • One primary metric chosen; guardrails listed (revenue per visitor, average order value, refund or cancellation rate)
  • Baseline conversion rate measured on the eligible audience over at least four weeks
  • MDE set as the smallest lift worth shipping, labeled relative or absolute
  • Sample size calculated at 95% confidence and 80% power
  • Duration rounded up to whole weeks: two-week floor, roughly eight-week ceiling
  • Launch window clear of promotions, holidays and planned site releases
  • Tracking verified on both variants, on mobile and desktop
  • Segments for analysis chosen in advance, two at most
  • Decision rule written for a positive, flat or negative interval
  • Stop date on the calendar, with no decisions before it
  • Mid-test checks limited to breakage: sample ratio, errors, tracking gaps

This is the testing discipline I set up in CRO and funnel work: fewer tests, sized properly, with results the finance team can plan around.

Get it built

If your testing tool keeps finding winners that never show up in revenue, the Growth Audit reviews your test history, tracking and backlog. It’s $1,500 fixed and credited if we continue. See pricing or get in touch.

FAQ

Frequently Asked Questions

How long should an A/B test run?

Until it reaches the sample size you calculated before launch, rounded up to whole weeks, with two weeks as a practical floor. If the math says more than about eight weeks, the test is too small for your traffic, so test a bolder change or use a different method.

Can I stop an A/B test early if it's already significant?

Not on a standard fixed-sample test. Significance swings as data comes in, and stopping at the first significant reading pushes the real false positive rate far above 5%. The exception is a tool built for sequential testing, where you follow its specific stopping rules.

What minimum detectable effect should I use?

The smallest lift that would be worth shipping and worth the traffic. Calculate the sample size it needs, and if the duration is unrealistic, raise the MDE and test a bigger change rather than loosening your significance standard.

What does 95% significance actually mean?

It means that if the variant truly made no difference, a gap at least this large would appear less than 5% of the time by chance. It does not mean there's a 95% chance the variant is better, and it says nothing about the size of the lift; the confidence interval does.

Work with me

Let’s find your biggest growth lever

Tell me about your growth challenge. I’ll tell you honestly if I can help — and if I can’t, who can.

  • ✓ No obligation
  • ✓ No sales script
  • ✓ Honest feedback
  • ✓ Clear next steps