A/B Test Significance Calculator
Visitors and conversions for control and variant. The tool runs a two-proportion z-test and tells you whether to trust the lift.
Verdict
–
Enter visitors and conversions for both arms
- Control rate
- –
- Variant rate
- –
- Relative lift
- –
- Variant vs control
- Absolute lift
- –
- Percentage points
- Z-score
- –
- P-value
- –
- Confidence (1 − p)
- –
- Not the chance B is better
- Z needed
- –
- To reach your confidence level
What significance does, and does not, tell you
A significant result means the gap you saw would be unusual if the two versions were really the same. That is all. The tool pools both arms into one conversion rate, works out how much two samples of this size would normally wander apart by chance (the standard error), and expresses your observed gap as a multiple of that wander (the z-score). A z-score near 2 or beyond is hard to get by luck alone, so the p-value drops below 0.05 and the result clears the 95% bar.
Significance is not the same as importance, and it is not a forecast. A 15% lift measured on 10,000 visitors per arm might really be anything from a few percent to well over 20%; the test tells you the direction is probably real, not that the number will hold. Small samples cut both ways: an underpowered test can miss a real winner and can also crown a fake one, which is why the tool warns when either arm has few conversions.
Two habits protect you. Decide the sample size before you start, using the sample size calculator, and do not stop early because the dashboard turned green. Checking repeatedly and stopping on a good day inflates false positives. And pick the test type before you look at the data. A one-tailed test asks only "is B better than A" and reaches significance more easily; a two-tailed test asks "is B different from A" and is the honest default when a variant could plausibly lose.
Formulas
- p₁
- = Control conversions ÷ Control visitors
- p₂
- = Variant conversions ÷ Variant visitors
- Pooled p
- = (Control conversions + Variant conversions) ÷ (Control visitors + Variant visitors)
- SE
- = √( Pooled p × (1 − Pooled p) × (1 ÷ n₁ + 1 ÷ n₂) )
- z
- = (p₂ − p₁) ÷ SE
- p-value (two-tailed)
- = 2 × (1 − Φ(|z|))
- p-value (one-tailed, B better)
- = 1 − Φ(z)
- Relative lift
- = (p₂ − p₁) ÷ p₁
- Absolute lift
- = p₂ − p₁
- Significant when
- p-value < 1 − Confidence level
Frequently asked questions
What does a p-value of 0.03 actually mean?
If the control and variant truly converted at the same rate, you would see a gap at least this large about 3% of the time purely from random variation in who happened to land in each group. It is not the probability that the variant is better, and it says nothing about the size of the real lift. A small p-value with a small sample can still come with a wide range of plausible true lifts, including some close to zero.
Should I use a one-tailed or two-tailed test?
Two-tailed is the safer default. It asks whether the variant is different in either direction, which is what you usually need to know, since a variant can lose. A one-tailed test asks only whether the variant is better, and it reaches significance with a smaller gap because all of the allowed error sits on one side. Use it only when you decided before the test that a loss and a tie would be treated identically, and say so when you report the result.
Can I stop the test as soon as it reaches 95%?
Not if you want the 95% to mean 95%. Checking daily and stopping on the first significant reading is called peeking, and it can push the real false-positive rate well above the 5% you think you are accepting, because a noisy test will cross the line at some point by chance. Decide the sample size in advance with the sample size calculator, run for whole weeks so weekday and weekend traffic are both included, and read the result once at the end.
More free tools
Browse all free tools →- Conversion Rate Calculator Calculate conversion rate from visitors and conversions, and see what a lift would be worth in extra conversions and revenue.
- A/B Test Sample Size Calculator Calculate how many visitors an A/B test needs per variant to detect a given lift at a chosen confidence and power.
- Funnel Conversion Calculator Model a multi-stage funnel, see where the biggest drop-off is, and test what fixing one stage does to the bottom line.
Draw it out before the board does.
Monday Brief plus a Thursday deep dive. Free, no paid tier.