Menu

A/B Test Sample Size Calculator

How many visitors each variant needs before the test can say anything, and how many days that is at your traffic. Set the smallest lift worth detecting, the confidence and the power, and read the count with the formula underneath.

By Nethanel Bar, Co-founder & CEO

Last updated

Ready to actually learn to code?

Coddy teaches you by writing real code in your browser - interactive lessons, instant feedback, and AI help when you get stuck.

Decide the sample size before you look at a single result

A sample size is a promise you make to yourself before the test starts: "I will look at the result when this many people have seen it, and not before." Without it, every A/B test drifts into the same story. The first week looks promising, the second week the lift shrinks, someone asks whether it is significant yet, and the test ends the day the bar turns green. This page gives you the number to write in the plan, so the end date is a fact rather than a mood.

The calculation is the standard two-proportion sample size, the one behind Evan Miller's calculator and Optimizely's, and it depends on four things: the baseline rate, the lift you want to be able to detect, the confidence level, and the power. Power is the one people skip. At 80% power, a real lift of the size you named is caught eight times in ten; a test with half the sample has far less than half the power, which is why an underpowered test is not "a rougher answer" but usually no answer at all.

The curve under the result is the part worth remembering. The visitors needed grow with the square of one over the lift: halve the lift you want to see and you need about four times the traffic. A site that can confirm a 30% lift in two weeks needs about nine times as long to confirm a 10% one. If the number of days comes out in months, that is the calculator telling you to test something bigger, or to move up the funnel to a metric more people reach, before it is telling you to wait.

What the sample size depends on

  • The minimum detectable effect is a decision, not a measurement. It is the smallest lift you would actually ship for. Setting it low to make the test "more sensitive" makes the test longer, not better.
  • Relative and absolute lifts are different animals. A 10% relative lift on a 5% baseline is half a percentage point. This page works in relative terms because that is how results are reported, and shows the target rate beside it so you can see both.
  • Extra variants cost more than their share. Testing A against B, C and D is three comparisons, and the confidence has to be split across them. The calculator applies the Bonferroni split and shows the alpha each comparison actually gets.
  • Days are a ceiling, not a floor. Even when the numbers say four days, run at least one full week, ideally two, because weekday and weekend visitors convert differently and a Tuesday-only test measures Tuesdays.
  • Power of 80% means one real winner in five is missed. If missing a winner is expensive for you, use 90% and accept the larger sample.

How to size an A/B test

  1. Enter the current conversion rate

    Use the rate of the page or step the test will change, measured over the same kind of visitors the test will see. Do not use a site-wide average when you are testing one funnel step.

  2. Choose the smallest lift worth detecting

    Ask what lift would make you ship the change. That is your minimum detectable effect. 10% relative is a common choice for a mature page; a redesign might justify 20% or more.

  3. Set confidence and power

    95% confidence and 80% power are the conventions. Raise confidence when a false winner is expensive to roll back; raise power when a missed winner is expensive to leave on the table.

  4. Add your daily visitors

    All variants together, counting only the visitors who will reach the test. The result turns into days and whole weeks.

  5. Write the end date down

    The calculator gives you the sample size per variant. Put it in the test plan with the date it implies, and do not read results before it. The peeking simulator tab shows why.

Sample size at a glance

Visitors per variant at 95% confidence and 80% power, two-sided. Multiply by two for a two-arm test.

Baseline rate+5% lift+10% lift+20% lift+50% lift
1%about 637,000about 163,000about 42,700about 7,800
3%about 208,000about 53,200about 13,900about 2,500
5%about 122,000about 31,200about 8,200about 1,500
10%about 57,800about 14,800about 3,800about 690
20%about 25,600about 6,500about 1,700about 290

Worked examples

A signup page at 5%, looking for +10%

Baseline 5%, minimum detectable effect 10% (target 5.5%), 95% confidence, 80% power, two-sided. Result: about 31,000 visitors per variant, 62,000 in total.

At 1,000 visitors a day that is 62 days, so nine whole weeks. Many teams would call that too long and either test a bolder change or accept that this page needs a lift of 20% or more to be testable in a fortnight.

A checkout step at 3%, looking for +20%

Baseline 3%, minimum detectable effect 20% (target 3.6%). Result: about 13,900 visitors per variant.

Lower baselines need more visitors for the same relative lift, because there are fewer conversions to count. A 20% relative lift here is only 0.6 percentage points, and the test has to tell 3.0% from 3.6%.

Three variants against a control

Baseline 5%, lift 10%, four variants including the control. The alpha per comparison becomes 0.05 / 3, and the per-variant sample grows by about a third; the total is four times the per-variant figure.

Adding variants is the fastest way to turn a two-week test into a two-month one. Test one strong idea against the control; run the next idea afterwards.

Sample size mistakes

  • Sizing the test after seeing early results. The lift you happen to observe on day three is not the minimum detectable effect; it is noise with an opinion.
  • Using an absolute lift where a relative one was meant. A "10% lift" on a 2% baseline is either 2.2% or 12%, and the sample sizes differ by a factor of hundreds.
  • Stopping at the sample size when the result is not significant and calling it a loss. Reaching the sample size means the test can detect the effect you named. It says nothing about smaller effects.
  • Counting page views as visitors. The formula assumes each unit is one independent visitor with one outcome; repeat views inflate the count and the confidence.
  • Ending on a partial week. If the sample size is reached on a Thursday, run to the following Sunday; weekday mix changes conversion rates more than most tested changes do.

Sample size FAQ

How long should I run an A/B test?
Until each variant has reached the sample size for the lift you want to detect, and for at least one full week, preferably two. Enter your daily visitors above and the calculator turns the sample size into days and whole weeks. Running longer than needed is a small waste; stopping early because the result looks good is how most false winners are declared.
What is a minimum detectable effect?
The smallest lift the test is designed to catch reliably, chosen before the test. It is usually expressed as a relative change, so a 10% MDE on a 4% baseline means the test can tell 4.0% from 4.4%. Smaller MDEs need dramatically more visitors: halving the MDE roughly quadruples the sample.
What does 80% power mean?
That if the true lift is exactly the size you named, a test this size will come out significant 80% of the time. The other 20% it misses a real winner. Power is the mirror image of confidence: confidence limits false alarms, power limits missed winners.
Why does a lower conversion rate need more visitors?
Because the test counts conversions, not visitors. At a 1% rate, 10,000 visitors give you 100 events to compare; at 10%, they give you 1,000. The same relative lift is far easier to see in the second case. This is why testing on a metric higher up the funnel, one more visitors reach, is often the only way a small site can test at all.
Can I run more than one variant?
Yes, and the calculator sizes it correctly, but each extra variant is another comparison against the control and the confidence has to be shared between them. Four variants at 95% without correction is a 14% chance of at least one false winner. The calculator splits alpha across the comparisons, which is why the per-variant number grows.
Does this work for revenue or average order value?
No, this page sizes a test on a conversion rate. Continuous metrics like revenue per visitor need the variance of the metric, which is usually far larger than a proportion's, and the sample sizes are correspondingly bigger. You need your own historical variance and a t-test based calculation for those.

Other developer tools

Coddy programming languages illustration

Learn to code with Coddy

GET STARTED