Sample Size Calculator

Sample size for an A/B test or a survey, with the relative and absolute effect spelled out and what halving the detectable effect costs.

Live output

Enable JavaScript to customise; default output below.

For

A test detects a difference between two groups. A survey estimates one number to a margin of error.

In survey mode this is the rate you expect; leave it at 50 if you have no idea, which gives the largest sample.

Significance level, %
Live preview sample-size.txt
Baseline conversion rate       3.2%
Improvement to detect          10% relative
  which is                     0.32% in absolute terms
  taking the rate to           3.52%
Significance level             95%, two sided
Power                          80%

Sample needed a group          49,773
Across 2 variants              99,546
Conversions a group, roughly   1,593

At this traffic
  visitors a day               4,200
  days needed                  23.7
  weeks                        3.4
  full weeks to run            4
  why whole weeks              behaviour differs by day, so a part-week biases the result

What a different target costs
  5% relative                  194,524 a group, 3.91× yours
  10% relative                 49,773 a group  ← yours
  20% relative                 13,012 a group, 0.26× yours
  40% relative                 3,534 a group, 0.07× yours

What more power costs
  80% power                    49,773 a group  ← yours
  85% power                    56,936 a group
  90% power                    66,635 a group
  95% power                    82,409 a group

The improvement is relative: 10% on a 3.2% baseline means reaching
3.52%, an absolute difference of 0.32%. Entering 10 when you meant 10
percentage points is the single most common mistake here, and the two
answers differ by orders of magnitude.

Sample size scales with the inverse square of the effect. Detecting half
the effect takes four times the data, which is why an ambitious test
target turns into a year of traffic, and why deciding what improvement
is worth detecting comes before any arithmetic.

Power of 80 percent means that if the effect is real and exactly the
size you specified, you will detect it 80 times in a hundred. Eighty
percent is the convention, and it means missing a real effect one time
in five, which is a worse trade than most people realise.

Stopping when the result looks significant invalidates it. Checking
repeatedly and stopping at the first p below 0.05 raises the false
positive rate far above 5 percent: with daily checks over a fortnight it
is above 20. Fix the sample size in advance, or use a sequential method
designed for peeking.

Testing several variants multiplies the chance of a false positive.
Three variants against a control is three comparisons, so a 5 percent
threshold gives about a 14 percent chance of at least one spurious
winner. Split the threshold across the comparisons, or accept the risk
knowingly.

Run for whole weeks. Behaviour differs by day, and a test ending on a
Tuesday afternoon has more Tuesdays in it than Sundays, which biases the
result in a way no sample size fixes.

The sample size assumes the observations are independent. A visitor who
sees the test on three devices counts three times, and a bot counts as
many times as it visits; both break the arithmetic before any of the
rest matters.

Output is valid and updates as you type.

“We want to detect a 10 percent improvement” and “we want to detect a 10 percentage point improvement” are different requests. On a 3.2 percent baseline the first means reaching 3.52 percent and needs about 50,000 visitors a group. The second means reaching 13.2 percent and needs about 400.

That mix-up is the most common error in sample size planning, so this states both every time.

The other thing worth knowing before the arithmetic: sample size scales with the inverse square of the effect. Halving the effect you want to detect quadruples the data you need, which is why an ambitious test target turns into a year of traffic.

How to use

  1. Pick a test or a survey. They are different calculations.
  2. For a test: the baseline rate, the relative improvement worth detecting, and the traffic.
  3. For a survey: the margin of error you want, and the population if it is small.

Example

A 3.2 percent baseline, detecting a 10 percent relative improvement:

Improvement to detect          10% relative
  which is                     0.32% in absolute terms
  taking the rate to           3.52%

Sample needed a group          49,773
Across 2 variants              99,546
Conversions a group, roughly   1,593

At this traffic
  visitors a day               4,200
  days needed                  23.7
  full weeks to run            4

What a different target costs
  5% relative                  194,524 a group, 3.91× yours
  10% relative                 49,773 a group  ← yours
  20% relative                 13,012 a group, 0.26× yours

What more power costs
  80% power                    49,773 a group  ← yours
  90% power                    66,635 a group

And a survey to ± 3 points at 95 percent confidence:

Responses needed             1,068
  at the worst case of 50%   1,068
  why 50 percent             p(1−p) is largest there, so it is the safe assumption

Pitfalls

Relative and absolute are not interchangeable. The two answers differ by orders of magnitude. Decide which you mean before touching a calculator, and say which one in the test plan.

Stopping early invalidates the result. Checking daily and stopping at the first p below 0.05 pushes the false positive rate well above 5 percent: over a fortnight of daily checks it is above 20. Fix the sample in advance, or use a sequential method built for peeking.

Power is the half everybody ignores. Eighty percent power means missing a real effect one time in five. That is the convention, and it is a worse trade than most people realise: a test that fails to detect a real 10 percent improvement has cost you the improvement as well as the traffic.

Several variants multiply the false positives. Three variants against a control is three comparisons, so a 5 percent threshold gives about a 14 percent chance of at least one spurious winner. Divide the threshold across the comparisons or accept the risk knowingly.

Run whole weeks. Behaviour differs by day, so a test that ends on a Tuesday afternoon has more Tuesdays than Sundays in it. No sample size fixes that bias.

Independence is assumed and often false. A visitor on three devices counts three times, a bot counts every visit, and a change that affects repeat visitors differently from new ones breaks the arithmetic before any of the rest matters.

In a survey, response bias beats sampling error. A 3 point margin of error on a 10 percent response rate is precise about the people who answered. They are not a random sample of the people who did not, and no sample size corrects for that.

Compatibility

Arithmetic in the browser: nothing is uploaded and nothing is stored.

The test formula is the standard two-proportion one: n = (z_α + z_β)² × (p₁(1−p₁) + p₂(1−p₂)) / (p₂ − p₁)², per group, with a two-sided significance level. The critical values come from a table of the handful of levels anybody uses rather than from an inverse-normal approximation, which keeps them exact to the digits people quote.

The survey formula is z²p(1−p)/e² with the finite population correction applied when a population is given. The published figure for ± 3 points at 95 percent, 1,068 responses, is fixed in the test suite, as is the way the sample quadruples when the detectable effect halves.

Both modes assume a binary outcome: converted or not, yes or no. For a continuous measure like revenue per visitor, the sample depends on the variance of that measure and these figures do not apply.

Frequently asked questions

How long should an A/B test run?
Until the planned sample is reached, and in whole weeks, and for at least one full business cycle. Two weeks is a common minimum regardless of traffic, because a single week can be unrepresentative.
What if I do not have the traffic?
Then test bigger changes. A test that needs a year is not a test; the options are to aim at a larger effect, to test on a higher-traffic page, or to accept a lower confidence level and treat the result as weak evidence.
Is 95 percent confidence required?
It is a convention, not a law. For a reversible change with low cost, 90 percent may be a perfectly reasonable threshold, and saying so in advance is what keeps it honest.
Why is my calculator’s number different?
Usually the pooled against unpooled variance, a one-sided against two-sided test, or relative against absolute effect. The first two make a few percent of difference; the third makes a hundredfold.
How many responses do I need for a survey?
About 400 for ± 5 points and about 1,070 for ± 3, at 95 percent confidence, almost regardless of population size once the population is above about 20,000. That last part surprises people and is correct.
Weekly drops

New tools, when there are new tools

One email when something worth using ships. No schedule to fill, so no filler.

Your address goes nowhere else, and one click unsubscribes.