Two thousand tests where A and B are the same page, run on your traffic. Watch how many a daily peeker would have called a winner, and how many the same tests fool someone who waits for the end date.
The most expensive habit in A/B testing, simulated on your numbers
Here is the habit. You start a test, and every morning you open the dashboard. The bar creeps up, dips, creeps up again, and one Tuesday it crosses 95%. You call it, ship the winner, and move on. The problem is that a p-value wanders. On a test where the two variants are identical, the p-value still crosses 0.05 on some days by chance; it just crosses back later. Stopping at the first crossing means you catch it on the way through, and the confidence level you thought you had is gone.
The size of the effect surprises people. At 95% confidence, a test read once at the end fools you about one time in twenty, which is the deal you signed. Read the same test every day for four weeks and stop at the first green day, and the simulator typically shows a false winner in a quarter or more of tests where nothing changed. Evan Miller's original write-up put it at over 20% for a daily peeker; Optimizely rebuilt its statistics engine in 2015 because its customers' false positive rate had climbed past that. Checking weekly instead of daily roughly halves the damage but does not remove it.
The simulation here is honest about what it is: each test draws conversions at the true rate for both arms, runs the ordinary two-proportion z-test at every look on everything collected so far, and records where a peeker would stop. Every path in the chart is one test; every dot is a morning someone would have announced a result. Set the real lift to +10% to see the other half of the story, where the peeker catches a real effect early but with an exaggerated lift, and would have "won" just as eagerly had the effect been zero.
What the simulation shows
The p-value is a random walk. Under the null it is uniformly distributed at any single look, which means that on any one morning there is a 5% chance it sits under 0.05. Over twenty-eight mornings those chances add up.
The false positive rate depends on how many times you look, not on how long the test runs. A four-week test read weekly is fooled less often than a two-week test read daily.
Early stops overstate the lift. When a peeker stops on a lucky day, the observed lift on that day is by definition unusually large. Shipped winners that "faded" after launch were often this.
The fix is either discipline or a different test. Discipline: fix the sample size and end date in advance and read once. A different test: sequential methods, like the always-valid p-values Optimizely, Statsig and Eppo use, are built to be read continuously at a cost in sample size.
Bayesian dashboards are not immune. A probability-to-be-best that is checked daily and acted on at a threshold has the same problem in a different costume.
How to use the simulator
1
Enter your conversion rate and daily visitors
The simulation draws conversions at your real rate, so the wandering of the p-value matches the noise you would actually see.
2
Set the planned length and how often you look
Daily, every three days or weekly. The gap between the peeker's rate and the honest rate is the price of that habit.
3
Keep the real lift at none
That makes every test an A/A test: any winner is a false alarm by construction. Switch to +10% or +30% afterwards to see the peeker on a real effect.
4
Read the two big numbers
The peeker's rate is how often the daily checker declared a winner. The end-date rate should be close to the alpha you chose; if it is, the honest test is doing what it promised.
5
Run it again
Every run draws new random tests. The rates move a little; the gap does not.
Typical false positive rates on A/A tests
At 95% confidence, one honest read at the end is fooled about 5% of the time. Rough figures from the simulation and the literature for a peeker who stops at the first significant look.
How often you look
Looks in 4 weeks
False winners
Once, at the end
1
about 5%
Weekly
4
about 10 to 13%
Every 3 days
9 to 10
about 15 to 20%
Daily
28
about 25 to 30%
Daily for 12 weeks
84
over 35%
Three habits, simulated
The daily checker
5% conversion, 1,000 visitors a day, 28 days, checked every morning. In our run: the peeker declared a winner in about 28% of 2,000 A/A tests; reading once at the end, 4.5%.
Six times the false alarm rate, for a test that felt careful because someone was watching it every day. Half those "winners" crowned B and half crowned A, because there was nothing to find.
The weekly checker
Same traffic, same length, checked once a week. In our run: about 12% for the peeker against 5% at the end.
Looking four times instead of twenty-eight helps a lot and still more than doubles the promised rate. There is no number of looks except one that keeps the promise.
A real +10% lift
Set the real lift of B to +10% on the same traffic. The honest test at 28 days has limited power for a lift that small at this traffic; the peeker stops early more often, on the lucky days.
This is the seductive case: the peeker seems to find things faster. But the days it stops on are the days the lift looked biggest, so the shipped estimate is inflated, and the same habit would have found a "winner" with the lift set to zero.
Peeking mistakes
Stopping at the first significant morning. This is the whole page. The confidence level only means what it says if you read the result once.
Extending a test that is "almost there". Deciding to run longer because the result is nearly significant is peeking in reverse, and inflates the false positive rate the same way.
Checking daily "just to make sure nothing is broken". Watching for bugs is fine; watching the p-value is not. Look at the traffic split and error rates, not the lift, until the end date.
Believing a Bayesian dashboard makes peeking safe. Acting on a probability threshold that you check daily has the same inflation.
Reporting the lift from the day you stopped. If you must stop early, report the interval, and expect the true lift to be smaller than the point estimate.
Peeking FAQ
What is the peeking problem in A/B testing?
Reading a fixed-horizon test repeatedly and stopping the first time it looks significant. Because the p-value fluctuates while data comes in, each look is another chance to cross the threshold by luck. A test read daily for a month at 95% confidence produces a false winner roughly a quarter of the time when nothing changed, instead of the 5% the confidence level promises.
Is it wrong to look at my test before it ends?
Looking is fine; acting is the problem. Check that traffic is splitting evenly, that both variants load, and that conversions are being recorded. Do not read the lift or the confidence, and do not let what you saw change the end date in either direction.
How do I A/B test if I want to check continuously?
Use a sequential method. Always-valid p-values, group sequential designs and the like are built to be read at any time; they pay for it with a larger sample size for the same power. Most modern platforms offer one. This simulator models the classic fixed-horizon test because that is what most spreadsheet and dashboard calculations, including this site's other tabs, are.
Does peeking matter if I check weekly?
Less, but yes. Four looks over four weeks at 95% typically double the false positive rate. The only number of looks that keeps the promised rate is one.
Why does the simulator use A/A tests?
Because with identical arms, every winner is a false alarm by construction, so the rate the simulator reports is exactly the false positive rate of the habit being simulated. Switching on a real lift shows the other side, where the peeker stops early on lucky days and reports an inflated effect.
Does a Bayesian test solve the peeking problem?
Not by itself. A Bayesian probability read continuously and acted on at a fixed threshold inflates false decisions the same way, as Georgi Georgiev and others have shown. Bayesian methods with a proper decision rule and a prior can behave well under continuous monitoring, but "Bayesian" alone is not the fix.