Free A/B Test Significance Calculator
Find out whether your A/B test winner is real or just noise - with confidence, p-value, and chance-to-beat - and plan the sample size and run time before you start.
Runs entirely on your device. Your files are processed in your browser and never uploaded. There is no server to send them to, so we never receive, store, or see your files - and filenames are hidden from our analytics too. Load this page, disconnect from the internet, and it still works.
Control (A)
Variant (B)
Statistically significant
You can be 95.8% confident the variant genuinely differs from control - above the 95% bar.
5.00%
Control rate
6.50%
Variant rate
+30.0%
Relative uplift
97.9%
Chance to beat control
Confidence: 95.84%
p-value: 0.0416
z-score: 2.04
Is your winning variant real, or just noise?
The most expensive mistake in conversion testing is calling a winner too early. Random variation alone can make one version look better for a while, and if you ship that "winner" you bank a lift that was never really there. This calculator settles it: enter the visitors and conversions for your control and variant, and it computes each conversion rate, the relative uplift, and - critically - whether the difference is statistically significant. It runs a two-proportion z-test for the p-value and confidence level, and adds a Bayesian-style "chance to beat control" that answers the plain-language question stakeholders actually ask: how likely is it that B really is better than A? Everything updates instantly in your browser as you type.
Plan the test before you run it
Significance after the fact is only half the job; the other half is designing the test so it can reach significance at all. The sample-size planner works backwards from what you want to detect. Give it your baseline conversion rate and the minimum detectable effect - the smallest lift worth catching - and it returns the number of visitors you need per variant at 95% confidence and 80% power, plus an estimated run time from your daily traffic. This is what stops teams from running underpowered tests that could never have produced a clear answer, and from stopping too soon or dragging on too long. Deciding the sample size up front and holding to it is the single most important discipline in trustworthy testing.
Test honestly to keep the results trustworthy
Statistics only protect you if you follow the rules they assume. Do not peek at significance and stop the moment you cross 95% - repeatedly checking and stopping on a good result dramatically inflates your false-positive rate. Run to your pre-planned sample size, and cover at least one full business cycle (usually a week or more) so weekday, weekend, and payday behavior are all represented. Watch for a minimum detectable effect you can actually achieve with your traffic - detecting tiny changes on low-traffic pages can require impractically large samples. Used this way, the calculator keeps your optimization program grounded in real, repeatable lifts rather than lucky noise.
Frequently Asked Questions
It means the difference between your variant and control is unlikely to be due to random chance. At 95% confidence, there is roughly a 5% probability of seeing a difference this large if the two versions were truly identical. Significance says the effect is probably real - it does not tell you the effect is large, will last, or is worth shipping; those are separate judgments about effect size and business value.
It is a Bayesian-flavored probability that the variant's true conversion rate is higher than the control's, estimated from the observed data. It answers the intuitive question - "how likely is B actually better than A?" - more directly than a p-value. A chance to beat of 97% means it is very likely the variant genuinely wins; around 50% means the data are essentially a coin flip so far.
It depends on your baseline conversion rate and the size of the change you want to detect - smaller effects and lower baselines need far more traffic. Use the sample-size planner: for example, detecting a 20% relative lift on a 5% baseline needs roughly a few thousand visitors per variant. As a rule, decide this number before you launch and do not stop the test until you reach it.
Because significance fluctuates as data accumulates, and if you keep checking and stop the instant it crosses 95%, you will frequently stop on a random high point - a practice called "peeking" that can push your real false-positive rate well above 5%. Decide the sample size in advance and evaluate once at the end (or use a proper sequential-testing method). This calculator supports that discipline by letting you plan the sample size first.
This calculator compares two groups at a time (control vs one variant), which is the classic A/B setup. For an A/B/n test with several variants, compare each variant against the control separately, but be aware that testing many variants raises the chance of a false positive somewhere - you may want to apply a stricter significance threshold (a multiple-comparisons correction) as the number of variants grows.
Turning Tests Into Compounding Growth
A calculator confirms a winner; a program produces a pipeline of them. We run structured experimentation and CRO that lifts conversion month over month.