TOOL 07 / TEST

← All tools

A/B Test Calculator

Plan the sample size before launch. Check significance after. Both halves of testing, one instrument.

%
%
The smallest lift worth detecting
/day
Sample per variant visitors needed
Total sample both variants
Conversions per variant expected at baseline
Days to run at your traffic
Duration vs a 30-day test budget

Already ran the test?

Check if your result is significant

Raw numbers only: visitors and conversions per variant. A is your control, B is the change.

CVR A control
CVR B variant
Relative uplift 95% CI —
P-value two-tailed
Confidence

How the sample size formula works

The calculator uses the standard approximation for a two-variant test at 95% significance and 80% power:

n = 16 × p(1−p) ÷ δ²

where p is your baseline conversion rate as a decimal and δ is the absolute lift you want to detect, so δ = p × MDE. With the defaults: p is 0.02, a 20% relative lift makes δ = 0.004, and the formula gives 19,600 visitors per variant. At 500 daily visitors per variant, that is a 40 day test. The duration falls straight out of arithmetic you can do before you build anything.

Checking finished results: the z-test

z = (p₂ − p₁) ÷ √( p̄(1−p̄)(1/n₁ + 1/n₂) )

For a finished test the calculator pools both variants into one blended rate , asks how large the observed gap is compared to the noise you would expect at these sample sizes, and converts that z-score into a two-tailed p-value. Below 0.05 the result clears the conventional 95% bar. The uplift readout carries a 95% confidence interval: the range the true lift plausibly sits in. A significant test whose interval scrapes zero is a real but possibly tiny win, so read the interval, not just the verdict.

What significance and power actually mean

95% significance means that if there is truly no difference between the variants, you will wrongly declare a winner about 1 time in 20. 80% power means that if the lift is real and as big as your MDE, the test will catch it 4 times out of 5, and miss it once. Neither is a guarantee. They are error rates you have agreed to live with, and they only hold if you fix the sample size in advance and read the result once, at the end.

The low-traffic trap

Sample size scales with the inverse square of the effect. Detect a lift half as small and you need four times the traffic. This is why small stores copying the testing culture of large ones end up with tests that never finish: a 5% MDE at a 2% baseline needs over 300,000 visitors per variant. If the days output here comes back in months, the answer is not more patience. Raise the MDE and test changes big enough to clear it.

Fewer, bolder tests

On modest traffic, your testing calendar should hold a handful of big swings per quarter: a different offer, a rebuilt hero, a new pricing frame. Each one either moves the number visibly or teaches you something real. Ten micro-tests that all end inconclusive teach you nothing except that your traffic cannot support micro-tests. Match the ambition of the change to the traffic you actually have.

FAQ

How long should an A/B test run?
Until it reaches the required sample size, and never less than one full week even if the sample allows it. Traffic behaves differently on a Tuesday than on a Saturday, so a test that only sees part of the weekly cycle is measuring the calendar, not the variant. Two full weeks is the practical minimum for most sites. Round your duration up to whole weeks.
What sample size do I need for an A/B test?
It depends on your baseline conversion rate and the smallest lift you want to detect. At a 2% baseline, detecting a 20% relative lift takes about 19,600 visitors per variant at standard settings. Halve the effect you want to detect and the required sample roughly quadruples, which is why chasing small lifts on modest traffic rarely works.
Why can't I stop the test the moment it hits significance?
Because checking repeatedly and stopping on the first significant reading inflates your false positive rate badly. A test peeked at daily can show significance at some point even when there is no real difference. Decide the sample size and duration before launch, then let it run. The math only holds if you evaluate once, at the end.
How do I choose the minimum detectable effect?
Set it to the smallest lift that would actually change a decision, then check whether your traffic can afford it. On low traffic, a 10% MDE produces tests that run for months, so test bigger swings: new headlines, new offers, new page layouts that could plausibly move things 20% or more. Small polish tests are a luxury of high-traffic sites.
What does the p-value actually mean?
It is the probability of seeing a gap at least this large between the variants if they were truly identical. A p-value of 0.03 means: if the change did nothing, only 3 tests in 100 would show a difference this big by pure luck. It is not the probability that B is better, and it says nothing about how big the win is. Pair it with the uplift and its confidence interval before acting.
Do these rules apply to ad creative tests?
Only loosely. Ad platforms do not split traffic evenly, they optimize delivery toward whichever variant looks better early, so the clean statistics here never fully apply. Treat creative tests as directional: give each variant enough spend to exit learning, compare on your primary metric, and accept more uncertainty than a proper landing page test carries.

Want the thinking behind the numbers?

One paid-ads breakdown a week. Free.

The kind agencies charge for: teardowns, setups, frameworks. Built for ecommerce brands that want to scale.

Get the free breakdown