Skip to main content

A/B Test Sample Size Calculator

Baseline conversion rate, the smallest lift worth detecting, and your confidence and power settings. Get the sample size per variant and the test duration at your traffic.

Free toolsConversion & TestingReviewed September 2026

Test settings

Current conversion rate of the page or flow you are testing.

Relative lift. 10% means 4.0% rising to 4.4%, not 14%.

Two-tailed. Limits false positives.

Limits missed real wins.

DurationHow long the test runs at your traffic.

Across all variants, for the page or flow being tested.

Including control. Traffic is split evenly.

Sample size per variant

–

Visitors each arm needs

Total sample
–
All variants
Test duration
–
At your daily traffic
Rate to detect
–
Baseline × (1 + MDE)
Absolute lift
–
Percentage points

Runs entirely in your browser. Nothing you enter is stored or sent anywhere. Last reviewed September 2026.

Sample size by minimum detectable effect, at the settings above

MDE (relative) Rate to detect Per variant Total Days

The highlighted row matches your MDE. Halving the MDE roughly quadruples the sample.

Why halving the lift you want to detect quadruples the sample

Sample size grows with the square of one over the effect you are trying to see. The noise in a measured conversion rate shrinks with the square root of the visitors, so to halve the noise you need four times the traffic. Detecting a 5% lift instead of a 10% lift means resolving a gap half as wide, which needs noise half as large, which needs four times the sample. The table above shows that curve at your own baseline.

The baseline rate matters almost as much. A 10% relative lift on a 1% conversion rate is a gap of 0.1 points, buried in far more noise than the 0.4 point gap the same relative lift produces at 4%. Low-conversion pages need much more traffic to test, which is why teams often test on a higher-volume upstream metric (add to cart, form start) and confirm the downstream effect separately.

Confidence and power are the two error rates you are willing to accept: confidence caps how often you call a winner that is not real, power caps how often you miss one that is. Tightening either raises the sample. Pick them once as policy rather than per test, then use the duration figure honestly: run whole weeks even if the sample fills early, and if the duration runs past a couple of months, raise the MDE or test something with more traffic instead of hoping.

Formulas

p₁
= Baseline rate
p₂
= p₁ × (1 + MDE)
p̄
= (p₁ + p₂) ÷ 2
z(α/2)
= Φ⁻¹(1 − α ÷ 2)
z(β)
= Φ⁻¹(Power)
n per variant
= ( z(α/2) × √(2 p̄ (1 − p̄)) + z(β) × √(p₁(1 − p₁) + p₂(1 − p₂)) )² ÷ (p₂ − p₁)²
Total sample
= n × Number of variants
Days
= Total sample ÷ Daily visitors, rounded up

Frequently asked questions

What minimum detectable effect should I use?

Start from the decision, not the statistics. Ask what lift would actually change what you do: ship the variant, roll it out to other pages, fund more work in that direction. If a 5% lift would not be worth the engineering cost of keeping it, there is no point sizing a test to detect 5%. Most teams land between 5% and 20% relative. Smaller than that and the required traffic grows fast; larger and you will only ever catch big, obvious wins.

What is statistical power and why is 80% the default?

Power is the probability that the test detects a lift of the chosen size when that lift is real. At 80% power, one real winner in five will be missed and read as "not significant." Raising power to 90% cuts that miss rate in half but adds roughly a third more visitors per variant. 80% is a convention that balances the cost of traffic against the cost of missed wins; use 90% when the decision is expensive to get wrong.

Can I shorten the test by adding more variants?

No. Each variant needs the full per-variant sample, so a test with three variants needs about 50% more total traffic than a two-arm test and takes 50% longer at the same daily volume. Extra variants also multiply the comparisons, which raises the chance that at least one appears significant by luck. If you must test several ideas, test the two you believe in most, or accept a longer run.

More free tools

Browse all free tools →