TOOL 07 / TEST
← All toolsA/B Test Calculator
Plan the sample size before launch. Check significance after. Both halves of testing, one instrument.
Already ran the test?
Check if your result is significant
Raw numbers only: visitors and conversions per variant. A is your control, B is the change.
How the sample size formula works
The calculator uses the standard approximation for a two-variant test at 95% significance and 80% power:
n = 16 × p(1−p) ÷ δ²
where p is your baseline conversion rate as a decimal and δ is the absolute lift you want to detect, so δ = p × MDE. With the defaults: p is 0.02, a 20% relative lift makes δ = 0.004, and the formula gives 19,600 visitors per variant. At 500 daily visitors per variant, that is a 40 day test. The duration falls straight out of arithmetic you can do before you build anything.
Checking finished results: the z-test
z = (p₂ − p₁) ÷ √( p̄(1−p̄)(1/n₁ + 1/n₂) )
For a finished test the calculator pools both variants into one blended rate p̄, asks how large the observed gap is compared to the noise you would expect at these sample sizes, and converts that z-score into a two-tailed p-value. Below 0.05 the result clears the conventional 95% bar. The uplift readout carries a 95% confidence interval: the range the true lift plausibly sits in. A significant test whose interval scrapes zero is a real but possibly tiny win, so read the interval, not just the verdict.
What significance and power actually mean
95% significance means that if there is truly no difference between the variants, you will wrongly declare a winner about 1 time in 20. 80% power means that if the lift is real and as big as your MDE, the test will catch it 4 times out of 5, and miss it once. Neither is a guarantee. They are error rates you have agreed to live with, and they only hold if you fix the sample size in advance and read the result once, at the end.
The low-traffic trap
Sample size scales with the inverse square of the effect. Detect a lift half as small and you need four times the traffic. This is why small stores copying the testing culture of large ones end up with tests that never finish: a 5% MDE at a 2% baseline needs over 300,000 visitors per variant. If the days output here comes back in months, the answer is not more patience. Raise the MDE and test changes big enough to clear it.
Fewer, bolder tests
On modest traffic, your testing calendar should hold a handful of big swings per quarter: a different offer, a rebuilt hero, a new pricing frame. Each one either moves the number visibly or teaches you something real. Ten micro-tests that all end inconclusive teach you nothing except that your traffic cannot support micro-tests. Match the ambition of the change to the traffic you actually have.
FAQ
How long should an A/B test run?
What sample size do I need for an A/B test?
Why can't I stop the test the moment it hits significance?
How do I choose the minimum detectable effect?
What does the p-value actually mean?
Do these rules apply to ad creative tests?
Related tools
Want the thinking behind the numbers?
One paid-ads breakdown a week. Free.
The kind agencies charge for: teardowns, setups, frameworks. Built for ecommerce brands that want to scale.