Sample Size Calculator
Sample size for an A/B test or a survey, with the relative and absolute effect spelled out and what halving the detectable effect costs.
Baseline conversion rate 3.2%
Improvement to detect 10% relative
which is 0.32% in absolute terms
taking the rate to 3.52%
Significance level 95%, two sided
Power 80%
Sample needed a group 49,773
Across 2 variants 99,546
Conversions a group, roughly 1,593
At this traffic
visitors a day 4,200
days needed 23.7
weeks 3.4
full weeks to run 4
why whole weeks behaviour differs by day, so a part-week biases the result
What a different target costs
5% relative 194,524 a group, 3.91× yours
10% relative 49,773 a group ← yours
20% relative 13,012 a group, 0.26× yours
40% relative 3,534 a group, 0.07× yours
What more power costs
80% power 49,773 a group ← yours
85% power 56,936 a group
90% power 66,635 a group
95% power 82,409 a group
The improvement is relative: 10% on a 3.2% baseline means reaching
3.52%, an absolute difference of 0.32%. Entering 10 when you meant 10
percentage points is the single most common mistake here, and the two
answers differ by orders of magnitude.
Sample size scales with the inverse square of the effect. Detecting half
the effect takes four times the data, which is why an ambitious test
target turns into a year of traffic, and why deciding what improvement
is worth detecting comes before any arithmetic.
Power of 80 percent means that if the effect is real and exactly the
size you specified, you will detect it 80 times in a hundred. Eighty
percent is the convention, and it means missing a real effect one time
in five, which is a worse trade than most people realise.
Stopping when the result looks significant invalidates it. Checking
repeatedly and stopping at the first p below 0.05 raises the false
positive rate far above 5 percent: with daily checks over a fortnight it
is above 20. Fix the sample size in advance, or use a sequential method
designed for peeking.
Testing several variants multiplies the chance of a false positive.
Three variants against a control is three comparisons, so a 5 percent
threshold gives about a 14 percent chance of at least one spurious
winner. Split the threshold across the comparisons, or accept the risk
knowingly.
Run for whole weeks. Behaviour differs by day, and a test ending on a
Tuesday afternoon has more Tuesdays in it than Sundays, which biases the
result in a way no sample size fixes.
The sample size assumes the observations are independent. A visitor who
sees the test on three devices counts three times, and a bot counts as
many times as it visits; both break the arithmetic before any of the
rest matters.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
“We want to detect a 10 percent improvement” and “we want to detect a 10 percentage point improvement” are different requests. On a 3.2 percent baseline the first means reaching 3.52 percent and needs about 50,000 visitors a group. The second means reaching 13.2 percent and needs about 400.
That mix-up is the most common error in sample size planning, so this states both every time.
The other thing worth knowing before the arithmetic: sample size scales with the inverse square of the effect. Halving the effect you want to detect quadruples the data you need, which is why an ambitious test target turns into a year of traffic.
How to use
- Pick a test or a survey. They are different calculations.
- For a test: the baseline rate, the relative improvement worth detecting, and the traffic.
- For a survey: the margin of error you want, and the population if it is small.
Example
A 3.2 percent baseline, detecting a 10 percent relative improvement:
Improvement to detect 10% relative
which is 0.32% in absolute terms
taking the rate to 3.52%
Sample needed a group 49,773
Across 2 variants 99,546
Conversions a group, roughly 1,593
At this traffic
visitors a day 4,200
days needed 23.7
full weeks to run 4
What a different target costs
5% relative 194,524 a group, 3.91× yours
10% relative 49,773 a group ← yours
20% relative 13,012 a group, 0.26× yours
What more power costs
80% power 49,773 a group ← yours
90% power 66,635 a group
And a survey to ± 3 points at 95 percent confidence:
Responses needed 1,068
at the worst case of 50% 1,068
why 50 percent p(1−p) is largest there, so it is the safe assumption
Pitfalls
Relative and absolute are not interchangeable. The two answers differ by orders of magnitude. Decide which you mean before touching a calculator, and say which one in the test plan.
Stopping early invalidates the result. Checking daily and stopping at the first p below 0.05 pushes the false positive rate well above 5 percent: over a fortnight of daily checks it is above 20. Fix the sample in advance, or use a sequential method built for peeking.
Power is the half everybody ignores. Eighty percent power means missing a real effect one time in five. That is the convention, and it is a worse trade than most people realise: a test that fails to detect a real 10 percent improvement has cost you the improvement as well as the traffic.
Several variants multiply the false positives. Three variants against a control is three comparisons, so a 5 percent threshold gives about a 14 percent chance of at least one spurious winner. Divide the threshold across the comparisons or accept the risk knowingly.
Run whole weeks. Behaviour differs by day, so a test that ends on a Tuesday afternoon has more Tuesdays than Sundays in it. No sample size fixes that bias.
Independence is assumed and often false. A visitor on three devices counts three times, a bot counts every visit, and a change that affects repeat visitors differently from new ones breaks the arithmetic before any of the rest matters.
In a survey, response bias beats sampling error. A 3 point margin of error on a 10 percent response rate is precise about the people who answered. They are not a random sample of the people who did not, and no sample size corrects for that.
Compatibility
Arithmetic in the browser: nothing is uploaded and nothing is stored.
The test formula is the standard two-proportion one:
n = (z_α + z_β)² × (p₁(1−p₁) + p₂(1−p₂)) / (p₂ − p₁)², per group, with a two-sided significance level.
The critical values come from a table of the handful of levels anybody uses rather than from an
inverse-normal approximation, which keeps them exact to the digits people quote.
The survey formula is z²p(1−p)/e² with the finite population correction applied when a population is
given. The published figure for ± 3 points at 95 percent, 1,068 responses, is fixed in the test suite,
as is the way the sample quadruples when the detectable effect halves.
Both modes assume a binary outcome: converted or not, yes or no. For a continuous measure like revenue per visitor, the sample depends on the variance of that measure and these figures do not apply.