Menu

A/B Test Calculator

Type the visitors and conversions of each variant. You get the verdict in one sentence, the p-value with its working, the interval, the Bayesian chance B is better, a check that the traffic split is not broken, and the smallest lift this test could ever have seen.

By Nethanel Bar, Co-founder & CEO

Last updated

Ready to actually learn to code?

Coddy teaches you by writing real code in your browser - interactive lessons, instant feedback, and AI help when you get stuck.

What this calculator tells you that the others do not

Every A/B test calculator prints a p-value. Most stop there, and that is how founders end up staring at 87% confidence wondering what it means. This one answers the questions that actually decide what you do next. Is the difference bigger than chance would produce? How big could the true lift be, at worst and at best? If the test came out "not significant", could it ever have found anything at this sample size? And is the traffic split itself healthy, or are you comparing two groups that were never equal?

It runs the standard two-proportion z-test, the same one used by Optimizely's classic engine, VWO's frequentist mode and every statistics textbook, and shows each line of the arithmetic: the pooled rate, the standard error, the z-score and the tail area. Beside it sits the Bayesian answer, the probability that B's true rate is above A's, because "there is a 91% chance B is better" is the sentence most people were hoping the p-value would give them, and it is a different number.

We built it after looking at our own dashboard. Coddy has run about twenty named experiments, ten of them per-country pricing tests on small slices of traffic, and our internal tool showed a confidence bar with the words "Keep running" under 95%. That is optional stopping with a friendly face. The three tabs beside this one, sample size, the reality check and the peeking simulator, exist because a calculator that only reports significance is a calculator that helps you fool yourself faster.

What to know before you trust the number

  • "Not significant" does not mean "no difference". It means the data cannot tell. The tile that says the smallest lift this test could see is the honest reading: if your observed lift is below it, the test was never going to answer.
  • The p-value is not the chance you are wrong. It is the chance of seeing a gap this big if the two variants were identical. How likely the winner is real also depends on how often your ideas work, which the reality check tab turns into a number.
  • Stopping on the first day the bar turns green is the most expensive habit in A/B testing. Checking daily and stopping early turns a 5% false-positive rate into 20% or more. The peeking simulator shows it happening on your own traffic.
  • A broken split poisons everything. If you planned 50/50 and got 46/54, something is routing users before they reach the test: a caching layer, a bot filter, a redirect. The calculator runs that check first and refuses to give a verdict while it fails.
  • Under about 25 conversions per arm the normal approximation behind the z-test is a sketch, not a measurement. The calculator says so instead of printing three decimal places of false precision.

How to use the A/B test calculator

  1. Enter visitors and conversions for A and B

    Visitors are the people who were bucketed into that variant, not page views. Conversions are the ones who did the thing you are measuring: signed up, paid, clicked. Both must be counted the same way for both arms.

  2. Pick the confidence and the hypothesis

    95% two-sided is the default and the honest choice when B could be better or worse. One-sided is only for when a worse B would be ignored in exactly the same way as no change, which is rarer than people think.

  3. Read the verdict, then the interval

    The banner says what the numbers support. The relative lift tile shows the interval: the true effect probably sits anywhere inside it, and the low end is the number to plan with, not the point estimate.

  4. Check the two warnings

    If the split is flagged, fix the test before reading anything else. If the smallest detectable lift is bigger than your observed lift, the test was underpowered: go to the sample size tab and see how many visitors it would take.

  5. Copy the result or the link

    Copy result gives you the summary for a message or a document. Copy link puts the four numbers in the URL, so whoever opens it sees the same calculation.

Reading the outputs

What each tile means, in one line.

TileWhat it isHow to read it
Relative lift(rate B minus rate A) divided by rate AThe headline. The interval beside it is the range the true lift probably sits in.
p-valueChance of a gap this big if A and B were identicalUnder 0.05 at 95% confidence means significant. It is not the chance the result is wrong.
Confidence1 minus p, as a percentageThe number most tools print. 87% is not almost 95%; it is a coin the wrong side of the line.
Chance B is betterBayesian probability B's true rate is above A'sThe sentence people want. 91% means B is probably better, not that it is 91% better.
Smallest lift this test could seeThe relative lift detectable at 80% power with these arm sizesIf your observed lift is below it, "not significant" means "could not tell".
Traffic splitObserved share of visitors in each arm against the planFlagged when the split is off by more than chance allows. Fix the test before reading results.

Three tests, three different lessons

The classic almost-there

A: 10,000 visitors, 500 conversions (5.00%). B: 10,000 visitors, 560 conversions (5.60%). Relative lift +12%, p = 0.058.

Not significant at 95%, and the interval runs from about minus 0.4% to plus 24%. The honest reading is "probably a small positive, could be nothing". The smallest lift this test could detect is about 17%, so a 12% effect was always going to land in the grey zone. It needs roughly twice the traffic, not another week of hoping.

The broken split

A: 5,000 visitors, 100 conversions. B: 4,300 visitors, 120 conversions. Planned split 50/50.

B looks like a clear winner at +40%. The calculator refuses to say so: a 5,000 to 4,300 split has a chance of under one in a million of happening by accident at 50/50. Something is removing users from B before they are counted, often a redirect that fails for one browser, or a bot filter that treats the variant differently. Whatever it is, the two groups are not comparable and the lift is meaningless.

The small site

A: 400 visitors, 12 conversions (3.0%). B: 400 visitors, 19 conversions (4.75%). Relative lift +58%, p = 0.20.

A 58% lift sounds enormous, and the p-value is about 0.19. With 12 and 19 conversions the calculator flags low data: the numbers swing so much at this size that a single extra signup moves the rate by a quarter of a point. At this traffic the smallest lift the test could confirm in a month is over 100%. That is not a reason to run it longer; it is a reason to read the reality check tab.

Mistakes this calculator is built to catch

  • Reading 90% confidence as "probably fine". At 95% the line is 95%. The number under it means the data did not clear the bar you set, and moving the bar after looking is the mistake, not the number.
  • Calling a non-significant test a tie. A tie is a tiny interval around zero. A non-significant test with a wide interval is an unanswered question, and the sample size tab says what answering it would cost.
  • Declaring a winner the first morning the bar turns green. The peeking simulator runs two thousand tests where nothing changed and shows how many a daily peeker would have shipped.
  • Testing five variants and reporting the best one at 95%. Five comparisons at 5% each is a 23% chance of a spurious winner. The sample size tab splits alpha for you when you add variants.
  • Comparing page views to unique users, or counting conversions over a different window than visitors. The z-test assumes each visitor is one independent trial; anything else inflates the confidence.
  • Ignoring the split. A 48/52 split on 100,000 visitors is not chance; it is a bug, and every result downstream of it is untrustworthy.

A/B test calculator FAQ

How do I know if my A/B test is statistically significant?
Enter the visitors and conversions of each variant, pick a confidence level (95% is standard), and read the verdict. The test is significant when the p-value is below 1 minus your confidence, so under 0.05 at 95%. This page also tells you the interval for the true lift and whether the test could have detected the effect you are looking at, which a bare p-value cannot.
What does a p-value of 0.08 mean for my test?
That if A and B truly converted at the same rate, a gap at least this big would show up about 8% of the time by chance. It is not the probability the result is wrong, and it is not "92% confidence that B is better". At 95% confidence it means the test has not cleared the bar; look at the smallest detectable lift tile to see whether it ever could have.
Is 90% confidence good enough?
It can be, if you decided on 90% before the test started and you accept a one-in-ten false positive rate. What is never good enough is running at 95%, seeing 91%, and deciding 90% was the real threshold all along. The confidence level is a promise about how often you are willing to be fooled; changing it after looking breaks the promise.
How many conversions do I need per variant?
There is no magic number, but under about 25 conversions per arm the z-test is a rough sketch, and most real tests need hundreds. The number that matters is the lift you want to detect: at a 5% baseline, seeing a 10% relative lift at 95% confidence and 80% power takes about 31,000 visitors per variant, roughly 1,550 conversions each. The sample size tab computes it for your own rates.
What is the difference between the frequentist p-value and the Bayesian chance B is better?
The p-value asks "how surprising is this gap if nothing changed?" and gives a yes or no at a threshold. The Bayesian number asks "given what I saw, how likely is B's true rate to be above A's?" and gives a probability. On the same data they usually agree in direction: a p of 0.05 sits near a 97% chance B is better on a two-sided test. This page shows both because the Bayesian sentence is the one people mean when they say "confidence", and the p-value is the one their tool prints.
Why does the calculator refuse to give a verdict when the split is off?
Because a sample ratio mismatch means the two groups were not assigned at random, and randomisation is the whole reason an A/B test works. If you planned 50/50 and the observed split has a chance of under one in a thousand of arising by luck, some mechanism is deciding who lands in which arm, and it could be correlated with conversion. The lift you see might be that mechanism, not your change.
Can I use this calculator for revenue per visitor or average order value?
No. This page tests a proportion: each visitor either converted or did not. Revenue per visitor is a continuous, heavily skewed metric and needs a t-test or a bootstrap on the per-user amounts, which needs the raw data rather than four totals. The t-test calculator in the Coddy math calculators handles the comparison of means once you have the per-user values.
Where does the smallest detectable lift come from?
From the same formula as the sample size calculator, run backwards. Given your arm sizes and A's rate, it finds the relative lift that 80% of tests this size would catch at your confidence level. If the lift you observed is below it, your test was underpowered for that effect, and "not significant" carries almost no information about whether the change worked.

Other developer tools

Coddy programming languages illustration

Learn to code with Coddy

GET STARTED