How do I know if my A/B test is statistically significant?
Enter the visitors and conversions of each variant, pick a confidence level (95% is standard), and read the verdict. The test is significant when the p-value is below 1 minus your confidence, so under 0.05 at 95%. This page also tells you the interval for the true lift and whether the test could have detected the effect you are looking at, which a bare p-value cannot.
What does a p-value of 0.08 mean for my test?
That if A and B truly converted at the same rate, a gap at least this big would show up about 8% of the time by chance. It is not the probability the result is wrong, and it is not "92% confidence that B is better". At 95% confidence it means the test has not cleared the bar; look at the smallest detectable lift tile to see whether it ever could have.
Is 90% confidence good enough?
It can be, if you decided on 90% before the test started and you accept a one-in-ten false positive rate. What is never good enough is running at 95%, seeing 91%, and deciding 90% was the real threshold all along. The confidence level is a promise about how often you are willing to be fooled; changing it after looking breaks the promise.
How many conversions do I need per variant?
There is no magic number, but under about 25 conversions per arm the z-test is a rough sketch, and most real tests need hundreds. The number that matters is the lift you want to detect: at a 5% baseline, seeing a 10% relative lift at 95% confidence and 80% power takes about 31,000 visitors per variant, roughly 1,550 conversions each. The sample size tab computes it for your own rates.
What is the difference between the frequentist p-value and the Bayesian chance B is better?
The p-value asks "how surprising is this gap if nothing changed?" and gives a yes or no at a threshold. The Bayesian number asks "given what I saw, how likely is B's true rate to be above A's?" and gives a probability. On the same data they usually agree in direction: a p of 0.05 sits near a 97% chance B is better on a two-sided test. This page shows both because the Bayesian sentence is the one people mean when they say "confidence", and the p-value is the one their tool prints.
Why does the calculator refuse to give a verdict when the split is off?
Because a sample ratio mismatch means the two groups were not assigned at random, and randomisation is the whole reason an A/B test works. If you planned 50/50 and the observed split has a chance of under one in a thousand of arising by luck, some mechanism is deciding who lands in which arm, and it could be correlated with conversion. The lift you see might be that mechanism, not your change.
Can I use this calculator for revenue per visitor or average order value?
No. This page tests a proportion: each visitor either converted or did not. Revenue per visitor is a continuous, heavily skewed metric and needs a t-test or a bootstrap on the per-user amounts, which needs the raw data rather than four totals. The t-test calculator in the Coddy math calculators handles the comparison of means once you have the per-user values.
Where does the smallest detectable lift come from?
From the same formula as the sample size calculator, run backwards. Given your arm sizes and A's rate, it finds the relative lift that 80% of tests this size would catch at your confidence level. If the lift you observed is below it, your test was underpowered for that effect, and "not significant" carries almost no information about whether the change worked.