Skip to content
Beyond Vision

Free A/B Test Significance Calculator.

Put visitors and conversions for both variants in and this runs a two proportion z test, which is the standard way to ask whether a difference in conversion rate is bigger than chance. It also does two things most calculators skip: it corrects the threshold for how many times you have checked the running test, because every look is another chance to cross the line by luck, and it shows the range the real difference sits in rather than only a yes or no. A result whose range still crosses zero has not told you which variant is better.

Free, no sign up, no email. It runs in your browser, so nothing you type is sent to us or to anyone else.

The Two Variants

Confidence
Times You Have Checked

The Answer

Not proven yet

B is running at 5.80% against A’s 5.00%, which looks like +16.1%. It is not separable from chance yet.

Rate, A
5.00%
Rate, B
5.80%
Relative Change
+16.1%
P Value
0.0813
Confidence
91.9%
Still Needed, Each
7,626

The Range The Truth Sits In

The real difference is somewhere between -0.10% and +1.71% in percentage points of conversion rate. That range crosses zero, so the data is still consistent with B being worse than A. Whatever the headline number says, the direction is not settled.

A two proportion z test, two tailed, with a Sidak correction when more than one look is declared. It assumes visitors were split at random and counted once each. It cannot tell you whether the traffic was split fairly, whether a bot skewed one side, or whether the difference is worth shipping. Nothing you type here is sent anywhere.

How this works.

Checking a running test is how most false winners happen

Significance at 95% means that if there were no real difference, you would see a result this extreme about one time in twenty. That guarantee holds for one look, at the end. It does not survive being checked every morning.

Every look is another roll. Check a test five times and the chance of crossing the line somewhere by luck is closer to one in four than one in twenty. This is why tests that looked like clear winners on a Wednesday so often fail to reproduce: the winner was the peeking, not the variant. Say how many times you have checked and the threshold here is tightened to pay for it.

A p value is not the size of the win

The p value answers one narrow question: how surprising is this difference if nothing is really going on. It says nothing about how large the difference is, and a large sample will make a difference of a tenth of a percentage point statistically significant without it being worth anybody's afternoon.

The range underneath the headline is the more useful number. It says where the real difference plausibly sits. When it runs from plus one to plus nine percent, you have a win of unknown size. When it runs from minus two to plus eight, you do not yet know which variant is better, whatever the headline lift says.

Why it refuses to answer on small numbers

The z test rests on a normal approximation, and that approximation needs enough of both outcomes to describe the data. Below roughly five conversions and five non conversions on either side it stops holding, and the p value it produces is arithmetic rather than information.

So on numbers that small this says so instead of returning a confident figure. A calculator that answers anyway is the more common design and the less honest one.

Decide the sample size before you start, not after

The right way to run a test is to work out how many visitors you need before it starts, run to that number, and read it once. The figure here for what is still needed is calculated from the difference currently showing, which is a reasonable guide and not a promise: if the true difference is smaller than what you have seen so far, it will take longer than this says.

It is also the number that tells you when not to bother. If resolving the difference you care about needs forty thousand visitors per variant and the page gets nine hundred a week, that test is a year long and the honest answer is to make a bigger change rather than measure a small one.

What the arithmetic cannot see

Whether the split was fair. Whether a bot hit one variant. Whether a promotion ran during half the test. Whether the two groups were different kinds of people because the traffic source changed mid flight. All of those produce a difference in conversion rate, and the test reports them exactly as it would report a real effect.

And it cannot tell you whether the winning variant is worth shipping. A two percent lift on a page that takes a week to rebuild is a different decision from the same lift on a headline change, and no calculator has any view on that.

Questions we get asked.

How do I know if my A/B test is statistically significant?

Enter visitors and conversions for both variants. The test compares the two conversion rates and returns a p value: below your chosen threshold, usually 0.05 for 95% confidence, the difference is unlikely enough to be chance that it is worth acting on. Check the range underneath as well, because a significant result whose range crosses zero has not settled the direction.

Why does checking my test more often make it harder to reach significance?

Because every look is another opportunity to cross the line by luck. At 95% confidence a single look carries a one in twenty false positive rate, but five looks carry closer to one in four. Declaring how many times you have checked lets the threshold be tightened to keep the real error rate at the level you asked for.

How many visitors do I need for an A/B test?

It depends on your baseline conversion rate and the size of the difference you want to detect: smaller effects need far more traffic. This shows what would be needed to resolve the difference currently showing, which is a guide rather than a guarantee. Working it out before the test starts is better practice than watching and waiting.

What is a good confidence level for an A/B test?

95% is the convention and is usually right. Raise it to 99% when the change is expensive to build or hard to reverse. Lowering it to 90% is defensible for cheap, reversible changes where being wrong occasionally costs less than moving slowly.

My test says B is 15% better but not significant. What does that mean?

That the difference you are seeing is well within what chance produces at your current sample size. It does not mean B is not better; it means the data cannot yet tell. Keep running to the sample size you need, and read it once you get there.

Is anything I type here stored or sent anywhere?

No. It all runs in your browser. There is no server behind this tool, no sign up and no email gate.
Every free tool

Ready to move the numbers?

Let us talk about your current goals, what is working, what is not, and where the biggest opportunities may be.

hello@beyondvisiondigital.comMonday to Friday, we reply within one business day

Tell us what you are working on.

We will take a look and get back to you.