Free tool · statistical significance for split tests

Free A/B test significance calculator

An A/B test significance calculator tells you whether the difference between two versions of a page or ad is likely to be real or down to chance. Enter the visitors and conversions for each version, and this free calculator runs a two-proportion z-test and shows both conversion rates, the uplift and the p-value.

By the Dolphin Analytics team · Last updated · Test method from the NIST/SEMATECH e-Handbook of Statistical Methods

Confidence level
People or sessions that saw version A.
Of those, how many converted.
People or sessions that saw version B.
Of those, how many converted.
Result B wins: the uplift is significant z = 2.38, p = 0.0175, threshold 0.05
Likely conversion rate for each versionLikely conversion rate for each versionVersion A converts at 3.00%, with a 95% range of 2.67% to 3.33%. Version B converts at 3.60%, with a 95% range of 3.23% to 3.97%.2.5%3.0%3.5%4.0%A 3.00%B 3.60%
Version A converts at 3.00%, with a 95% range of 2.67% to 3.33%. Version B converts at 3.60%, with a 95% range of 3.23% to 3.97%. Each curve shows where that version's true rate probably sits. The less the curves overlap, the stronger the evidence that one version really converts better.
From the same numbers
Conversion rate A
3.00%
Conversion rate B
3.60%
Uplift (B against A)
+20.0%
p-value
0.0175
Significant when the p-value is below 1 minus the confidence level, for example 0.05 at 95%.

Runs in your browser. Nothing you type is sent to Dolphin Analytics.

What does statistically significant mean in an A/B test?

A result is statistically significant when a gap as large as the one you saw would rarely happen by chance if the two versions really performed the same. The p-value measures that chance. At 95% confidence the cut-off is 0.05, so a p-value below 0.05 counts as significant.

How does this calculator work?

It runs a two-sided, two-proportion z-test, the standard large-sample test for comparing two conversion rates, as set out in the NIST/SEMATECH e-Handbook of Statistical Methods. It pools both groups to estimate a shared conversion rate, measures how far apart the two rates are in standard errors (the z-score), and turns that into a p-value.

StepFormulaWorked example
Rate Aconversions ÷ visitors300 ÷ 10,000 = 3.00%
Rate Bconversions ÷ visitors360 ÷ 10,000 = 3.60%
Pooled rateall conversions ÷ all visitors660 ÷ 20,000 = 3.30%
Standard error√(pooled × (1 − pooled) × (1 ÷ visitors A + 1 ÷ visitors B))0.2526%
z-score(rate B − rate A) ÷ standard error0.60% ÷ 0.2526% = 2.38
p-valuetwo-sided, from the z-score0.0175, significant at 95%

When should I stop an A/B test?

Decide the sample size and run time before the test starts, and stop when you reach them. Checking every day and stopping the first time the result turns significant makes false winners much more likely, because a running test drifts above and below the cut-off by chance. Run for whole weeks so weekday and weekend visitors are both counted.

Frequently asked questions

Is this A/B test calculator free?

Yes. The Dolphin Analytics A/B test calculator is free, needs no sign-up and runs in your browser, so the numbers you enter never leave your device.

What confidence level should I use?

95% is the usual default. 90% gives faster answers but accepts more false winners; 99% is stricter and needs more traffic to reach a verdict.

What do the bell curves show?

Each curve shows where a version's true conversion rate probably sits, given the visitors you have tested. The peak is the rate you measured, and the width comes from the standard error: fewer visitors give a wider curve. Curves that barely overlap point to a real difference; curves that sit on top of each other mean the test needs more data.

Why does my testing tool show a different result?

Testing tools use different methods. Some use Bayesian statistics and report a probability to beat the control; others adjust for checking results many times. This calculator uses the classic fixed-sample z-test, so its answer can differ from tools that use those methods.

Can I use it for click-through rates?

Yes. It works for any test where each visitor either does or does not do one thing. Enter impressions as visitors and clicks as conversions.

How many visitors do I need?

It depends on your current conversion rate and the smallest uplift you care about. Small uplifts on low conversion rates need many thousands of visitors per version. When a group has fewer than 5 conversions or non-conversions, the calculator warns that the result is not reliable yet.

Related tools and guides

Talk to us

Data giving you a headache? We'll take it off your plate.

Tell us what's broken, or grab a time. Either way you hear from a person, not a sales script.

Send a message

We reply within one working day.

Add a few details (optional) The more we know up front, the faster we can tell you what's wrong and how to fix it.

Protected by an invisible spam check. Prefer email? olam@dolphinanalytics.co.uk

Calendly · 30 min

Book a call

Thirty minutes on Google Meet with the founder.

The booking lands on the same record as your message.