Menu

Can I A/B Test? The Traffic Reality Check

The question to ask before the first test, not after the tenth. Enter your monthly visitors, your conversion rate and the lift you would actually act on, and get a straight answer: go, stretch, or stop and do something else.

By Nethanel Bar, Co-founder & CEO

Last updated

Ready to actually learn to code?

Coddy teaches you by writing real code in your browser - interactive lessons, instant feedback, and AI help when you get stuck.

Most small sites cannot A/B test, and nobody tells them

Every A/B testing guide is written by a company with millions of visitors. Microsoft's experimentation team, Google, Booking.com and Airbnb run thousands of tests a year, and their advice is right for them: test everything, trust the data, ship what wins. A startup with 8,000 visitors a month reads the same advice, builds the same dashboard, and then cannot understand why every test ends at 87% confidence with a lift that shrinks every week. It is not doing it wrong. It is doing something that does not work at that size.

This page does the arithmetic the guides skip. At 8,000 visitors a month converting at 3%, confirming a 10% lift takes about 106,000 visitors: thirteen months of every visitor you have, with the page frozen the whole time. In four weeks the smallest lift you could confirm is about 40%. Those are not tuning problems. Ronny Kohavi, who ran experimentation at Microsoft and Airbnb, puts the practical floor for detecting single-digit lifts at a few hundred thousand users; below that, he suggests that A/B testing may be premature and user research the better spend.

We are writing this from experience. Coddy's own dashboard ran about twenty named experiments, ten of them pricing tests on single countries, each on a small slice of traffic, and reported "Keep running" whenever the confidence sat under 95%. Several were marked successful and shipped. Looking back with this calculator, some of those tests could not have detected the lifts they were credited with. The honest version of that dashboard is this page: what can this traffic answer, how long would it take, and how often would a winner be a false alarm?

What decides whether you can test

  • Conversions per month, not visitors, is the number that limits you. A page with 8,000 visitors at 3% has 240 conversions a month to split between two arms. The z-test needs hundreds per arm before it says anything.
  • The lift you would act on sets the bill. Be honest about it: if you would ship a change for a 3% lift, you need a test that can see 3%, and almost no small site can.
  • Time is a cost, not a fix. A test that takes a quarter freezes the page for a quarter, spans seasons and campaigns, and ends in a different business than it started in.
  • How often your ideas work changes what a winner means. If one idea in ten really moves the metric, then at 95% confidence and 80% power more than a third of your "significant" winners are false alarms. The base rate is not a technicality; it is most of the answer.
  • Below the line there are better methods. Talk to five users, watch session recordings, ship the change and compare before and after with wide error bars, or test on a metric more visitors reach, like a click rather than a purchase.

How to run the reality check

  1. Enter the monthly visitors who would see the test

    Not the whole site. If the test is on the pricing page, it is the pricing page's visitors. If it is a checkout step, it is the people who reach that step.

  2. Enter the current conversion rate of that step

    The calculator shows the conversions a month this implies, which is the real constraint.

  3. Name the lift you would act on

    The relative lift that would make you ship the change. Most people say 10% and mean it; be honest if yours is smaller.

  4. Read the verdict and the ladder

    Go means under a month. Stretch means under a quarter, with the page frozen the whole time. Stop means the test cannot answer the question at your traffic. The ladder shows what bigger lifts would cost instead.

  5. Pick a base rate and read the false positive risk

    Choose how often your ideas really work. The tile shows what share of significant results would be false alarms at that rate, which is the number to remember the next time a test comes out green.

Months to confirm a 10% lift at 95% confidence, 80% power

Two arms, every visitor entering the test. Read your row and your column.

Monthly visitors1% baseline3% baseline5% baseline10% baseline
3,000over 100 months35 months21 months9.8 months
8,00041 months13 months7.8 months3.7 months
25,00013 months4.3 months2.5 months5 weeks
100,0003.3 months5 weeks19 days9 days
500,00020 days6 days4 days2 days

Three sites, three verdicts

The early-stage SaaS

8,000 visitors a month, 3% signup rate, wants to see a 10% lift. Verdict: stop. About 13 months and 106,000 visitors. In four weeks the smallest confirmable lift is about 40%.

Two hundred and forty signups a month cannot resolve a 0.3 point difference. The team's time is better spent on the five-user interview and on changes big enough to show up without a test: a new headline is not one of them; a new pricing model might be.

The content site

50,000 visitors a month, 2% newsletter signup rate, wants a 10% lift. Verdict: stretch. About 3.2 months. In four weeks the smallest confirmable lift is about 19%.

Testable, but only for bold changes. Testing the form placement is fine if it might move signups by a fifth; testing button copy is not. Alternatively test on the click to the form, which far more visitors make.

The marketplace

400,000 visitors a month, 4% conversion, wants a 10% lift. Verdict: go. The visitors arrive in under a week; run two whole weeks anyway.

This is the size the A/B testing guides are written for. The remaining risk is the base rate: if one idea in five really works, roughly one significant winner in six is still a false alarm, and the peeking simulator tab shows how fast that gets worse with daily checks.

Low-traffic testing mistakes

  • Running the test anyway and reading it at 85%. The confidence level you set is a promise about how often you accept being fooled. Reading below it after the fact is the same as never having set one.
  • Testing tiny changes. Button colours and word swaps produce lifts of a few percent at best, which small sites cannot detect. If you must test, test something that could plausibly move the number by a third.
  • Testing at the bottom of the funnel. Purchases are rare; clicks are common. A test on the click that leads to the purchase has ten times the events and can finish.
  • Letting a test run for six months. Seasons, campaigns and product changes all leak into it, and the two arms end up measuring different eras.
  • Trusting a winner because it was significant. With few real winners among your ideas, a significant result is more often a false alarm than the p-value suggests. The false positive risk tile puts a number on it.
  • Treating "not significant" as "the idea failed". At low traffic almost nothing is significant. The idea might be fine; the test could not see it.

Reality check FAQ

How much traffic do I need to A/B test?
It depends on your conversion rate and the lift you want to see, which is why this page asks for both. As a rule of thumb, confirming a 10% relative lift on a 3% conversion rate takes about 100,000 visitors; confirming a 30% lift takes about 12,000. Sites under 10,000 visitors a month can usually only detect very large effects.
What should a small site do instead of A/B testing?
Talk to users, in person or on a call: five conversations find most of the problems on a page. Watch session recordings. Make bold changes, ship them, and compare a few weeks before and after, knowing the comparison is rough. Test on metrics more visitors reach. And when the traffic arrives, come back here and test properly.
What is false positive risk and why is it different from the p-value?
The p-value tells you how surprising the result would be if nothing changed. False positive risk tells you, of the tests that came out significant, how many had nothing behind them. It depends on how often your ideas really work: at a 10% base rate, 95% confidence and 80% power, about 36% of significant winners are false alarms. Large experimentation programs report that only 10 to 33% of ideas move their metric, which is why the base rate matters so much.
Is a 3.5 month test really impossible?
Not impossible, and the calculator says "stretch" rather than "stop" up to a quarter. But a page frozen for a quarter is a page nobody is improving, and a test that spans a season or a campaign compares two different periods. If the change is big enough to justify a quarter of waiting, it is probably big enough to ship and measure before and after.
Can I lower the confidence to 90% to make the test faster?
You can, if you decide it before the test and accept a one-in-ten false alarm rate on top of the base-rate effect. It shortens the test by roughly a third, not by the factor of ten a small site would need. Raising the lift you are testing for helps far more.
Why does the calculator say stop when the sample size calculator gives a number?
Both use the same formula. The sample size page tells you the count; this page divides it by your traffic and judges the result. A count of 106,000 visitors is a number; thirteen months is a verdict.

Other developer tools

Coddy programming languages illustration

Learn to code with Coddy

GET STARTED