A/B Test Significance Calculator
The two-proportion z-test, plus the parts that decide whether it means anything: sample size needed, the interval on the lift, and what peeking costs.
A: visitors 1,000
A: conversions 100
A: rate 10% (8.29% to 12.02%)
B: visitors 1,000
B: conversions 130
B: rate 13% (11.06% to 15.23%)
Absolute difference 3% points
Relative lift 30%
95% interval on the difference 0.207% to 5.793% points
z 2.103
p-value, two sided 0.0355
Threshold 0.050
Verdict significant
Visitors per variant this lift needs 888
Visitors per variant you have 1,000
p is 0.0355, below the 0.050 threshold, so the difference is
statistically significant at 95 percent. That means the data would be
this lopsided less than 5.0 percent of the time if the two variants were
identical. It does not mean the lift is 30%: the interval on the
difference runs from 0.207% to 5.793% points, and that range is the
honest answer.
Detecting a 30% relative lift at this baseline needs about 888 visitors
per variant, and you have 1,000. The test is adequately powered for the
effect it found.
A p-value is only valid for a test whose sample size was fixed in
advance. Watching a running test and stopping when it crosses the
threshold finds a winner in roughly a quarter of tests where nothing is
happening, because you gave yourself many chances at a 5 percent error.
Decide the sample size, run to it, then look.
Significance is not importance. On a large enough sample a 0.1 point
difference is significant and still not worth the engineering to ship
it. Work out what lift would pay for the change before the test, and
compare the interval against that rather than against zero.
This is the two-proportion z-test, which assumes each visitor is counted
once, assigned at random, and independent of the others. Sessions rather
than users, a redirect that leaks traffic, or a variant shown to
logged-in customers only all break that assumption in ways no arithmetic
can repair.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
The arithmetic is settled: a two-proportion z-test, pooled standard error, p-value off the normal distribution. Every calculator agrees on that part.
What most of them leave out is the part that decides whether the number means anything.
Peeking breaks it. A p-value is only valid for a test whose sample size was fixed in advance. Watching a running test and stopping when it crosses 95 percent finds a winner in roughly a quarter of tests where nothing is happening, because you gave yourself many chances at a one-in-twenty error.
The winner’s lift is overstated. If you stopped because the number looked good, you stopped on a high reading. Expect the real effect to be smaller than the one you measured.
Significance is not importance. A 0.1 point difference is significant on a million visitors and not worth shipping.
So this reports the sample size the test needed, the interval on the difference, and the correction for more than two variants, next to the p-value.
How to use
- Put in visitors and conversions for both arms. Count each visitor once.
- Set the confidence level and the number of variants including the control.
- Read the interval on the difference, not just the verdict.
Example
1,000 visitors each, 100 conversions against 130:
A: rate 10% (8.29% to 12.02%)
B: rate 13% (11.06% to 15.23%)
Absolute difference 3% points
Relative lift 30%
95% interval on the difference 0.207% to 5.793% points
z 2.103
p-value, two sided 0.0355
Threshold 0.050
Verdict significant
Visitors per variant this lift needs 888
Visitors per variant you have 1,000
The lift reads as 30 percent. The interval says the true difference is somewhere between 0.2 and 5.8 points, which on a 10 percent baseline is a lift between 2 percent and 58 percent. That range is the honest answer, and it is the number to take to a decision about whether to build the thing permanently.
Pitfalls
Do not stop early. Fix the sample size, run to it, then look. If you must monitor, use a sequential test designed for it, which pays for the extra looks by raising the bar.
Two variants is one comparison; four is three. Without a correction, the chance of at least one false winner rises with every arm: at 95 percent confidence and three comparisons it is about 14 percent rather than 5. The tool applies Bonferroni, which is blunt and defensible.
Sessions are not visitors. The test assumes each unit is counted once and assigned at random. Counting sessions counts a returning visitor twice, in whichever arm they land in, and understates the variance.
A redirect test leaks. If the variant is a different URL reached by a redirect, some of the traffic never arrives, and the ones that drop out are not random: they are the slow connections.
“Not significant” is not “no difference”. It means the data cannot rule out zero. The interval usually also cannot rule out an effect you would care about, which is a different statement from “the variants perform the same”.
One test at 95 percent confidence is wrong one time in twenty by design. Run twenty tests where nothing is happening and expect one winner. That is not a flaw in the method, it is the method.
Novelty wears off. A change that gets attention in week one can lose it by week four. That is a reason to run for whole business cycles rather than to a visitor count alone.
Compatibility
Everything runs in the browser: nothing is uploaded and nothing is stored.
The p-value uses an erf approximation accurate to about 1.5 × 10⁻⁷, which is four more digits than any decision needs, and it is checked in the test suite against the published critical values: 1.96 gives 0.05, 2.576 gives 0.01, 1.645 gives 0.10.
The interval on the difference uses the unpooled standard error, which is the correct one for an interval: pooling assumes the null hypothesis, and an interval is what you compute once you have stopped assuming it. That is why the interval can exclude zero in the same test where a pooled z sits just under the threshold, and the two are not in conflict.
Per-variant rate intervals are Wilson score intervals rather than the textbook normal approximation, which is too narrow at the small conversion rates real tests have.
The sample size figure is the standard two-proportion calculation at 95 percent confidence and 80 percent power. It is a planning number: trust the order of magnitude and not the last digit.
This is a frequentist test. A Bayesian approach answers a different and often more useful question, “what is the probability that B is better”, and it does not have the peeking problem in the same way. Neither is a substitute for a fixed plan.