Skip to the calculator

A/B Test Calculator

100% private

Tells you whether to ship it — not just a p-value. Your traffic and conversion numbers never leave this tab.

Where are you up to?
A · control B · variation Visitors Conversions

A conversion is whatever you are counting — a purchase, a sign-up, a click. Use the same definition on both sides, and count visitors, not visits.

Every calculation happens in this tab. Your traffic and conversion numbers are never sent anywhere — which matters more here than on most tools, because together they give away your revenue.

The verdict

Fill in visitors and conversions for both A and B.

Advanced options

How sure you want to be before calling a winner. 95% means you accept being wrong about one time in twenty.

Keep this two-sided unless you genuinely do not care about B being worse. One-sided reaches significance sooner, which is exactly why it is so often misused.

The chance of spotting a real improvement if there is one. Used for the “how much more traffic” answer and for planning.

This is what separates “call it a draw” from “keep testing”. Once we can rule out any change bigger than this, more traffic will not tell you anything new.

Checking usage…

The full numbers

What to watch out for

Step by step

How to work out if your A/B test result is real

Four steps, no account, no daily limit, and no sales call between you and the answer.

Pick what you need

Choose I have results if the test has been running and you need to decide what to do. Choose I’m planning one if you want to know how much traffic it will take before you start — which is the order it should really be done in.

Enter both variants

Visitors and conversions for A and for B. Visitors means unique people who saw each version, not page views; conversions means how many of them did the thing you care about. Use the same definition on both sides.

Read the verdict

You get one of four answers — ship B, keep A, keep testing or call it a draw — with the conversion rates, the lift and the odds that B is genuinely better beneath it.

Take it with you

One button copies the verdict and every number behind it as plain text, ready to paste into a ticket, a deck or a message to whoever has to sign the change off.

Private by design

Why your conversion numbers never leave this tab

Traffic volume and conversion rate, together, are close to a statement of your revenue. They are among the most commercially sensitive numbers a company has — and nearly every other A/B test calculator on the web is run by a conversion-optimisation vendor who would very much like to know them, and who also sells to your competitors. Typing them into a form on that vendor’s site is a bigger decision than it looks. Here there is no form and no server, and you can prove it in about ten seconds.

Nothing is uploaded

The calculator is downloaded to your browser once and then does every sum on your device. The numbers you type never enter a network request, so there is no copy of them for us to keep, lose, sell or be compelled to hand over.

Nothing is saved

We deliberately do not remember what you calculated, not even in your own browser’s storage. Close the tab and it is genuinely gone — there is no history of last quarter’s experiments sitting on a shared laptop.

No account, no sales call

No sign-up, no work email, no “book a demo to see your full results”, and no cap on how many tests you run. A calculator that runs on your own machine cannot meter you, and we are not trying to sell you testing software.

To be precise about what does leave your device — because “100% private” is easy to say and worth checking. Three things, none of them your numbers:

1. The first time you calculate something, this page sends one anonymous message to our counter saying “ab-test-calculator was used”. Not your visitors, not your conversions, not the verdict, nothing identifying you. It exists so the usage number in the panel is a real one rather than something we invented.
2. Google Analytics records the page view and that the tool was used, the same as on every other page of this site. It is told which of the two modes you used and which of the four verdicts came out — both picked from a fixed list, never a count, a rate or a lift. Block it and the calculator works exactly the same.
3. When the page first loads, the fonts it is set in are fetched from Google Fonts, which means Google sees your IP address at that moment — the same as on every other page of this site. Once the page has loaded, that request never happens again, which is why the offline test below works.

Your traffic and conversion figures are in none of them. They never enter a network request at all — and unlike a promise, that is something you can check. Turn off your wifi, reload this page and keep calculating. It still works. No server-side calculator can pass that test.

At a glance

Tool specifications

InputVisitors and conversions for two variants — or, in planning mode, a baseline rate and the improvement you want to detect
OutputA plain-English verdict, conversion rates, relative and absolute lift, confidence interval, p-value, z-score, the probability that B is better, and the sample size needed
Significance testTwo-proportion z-test on the pooled standard error, one- or two-sided
Confidence intervalNormal-approximation interval on the difference, using the unpooled standard error, at 90%, 95% or 99%
Probability B is betterBayesian, with uniform Beta(1,1) priors. Computed exactly up to 20,000 conversions in B, and by a normal approximation of the posteriors above that, where the two agree to better than one part in a million
Sample sizeStandard two-proportion formula with pooled variance under the null and unpooled under the alternative, no continuity correction. Reproduces Evan Miller’s published calculator
Split-ratio checkChi-square goodness of fit at p < 0.0005, to catch a broken test before you read its result
Variants supportedTwo — one control and one variation. Three or more needs a correction for multiple comparisons that this tool deliberately does not fake
Numerical accuracyFull double precision. The normal distribution uses Hart’s algorithm and its inverse uses Acklam’s with a Halley refinement, rather than the low-precision approximation most calculators ship
Libraries usedNone. No vendored bytes, no CDN, no third-party request
Rate limitNone. There is no server to rate-limit
Where processing happensIn your browser, on your device — nothing is uploaded
Saved anywhereNo. Not on a server, and not in your browser’s storage either
Works offlineYes, once the page has loaded, including as an installed app
Sign-upNot required, not offered
PriceFree

Questions

A/B testing, answered

Are my conversion numbers uploaded or stored anywhere?
No. The whole calculation runs inside your browser using code downloaded to your device, so the numbers you type are never placed in a network request and never reach us. We also do not save them to your browser’s storage, so they do not linger on a shared machine. This matters more here than on most tools: visitor counts and conversion rates together are effectively a revenue disclosure, and almost every other calculator of this kind is hosted by a testing vendor. The simplest proof: disconnect from the internet and keep calculating — it still works.
What does “statistically significant” actually mean?
It means the gap between your two variants is bigger than random chance comfortably explains. Suppose A and B were secretly identical: you would still not get exactly the same conversion rate from both, because different people saw them. The p-value is the answer to “if they really were identical, how often would I see a gap at least this big anyway?” At 95% confidence you are agreeing to act whenever that answer is under 5%. Note what it does not mean: it is not the probability that B is better, and a significant result is not the same as a result worth shipping.
My result isn’t significant. Should I keep running it or stop?
This is the question most calculators leave you holding, and it is why this one gives four answers rather than two. “Not significant” covers two completely different situations. Either you have not gathered enough traffic yet to see an effect that may well be there — keep testing, and the tool tells you roughly how many more visitors per variant it would take. Or you have gathered plenty, and the range of plausible differences has narrowed to something too small to care about — that is a draw, more traffic will not change it, and your time is better spent on a bolder change. The tool splits the two by checking whether the confidence interval fits inside the “smallest lift worth caring about” you set in the advanced options.
How many visitors do I need for an A/B test?
Far more than most people expect, and it depends almost entirely on two things: how often you convert now, and how small an improvement you insist on detecting. A site converting at 3% that wants to catch a 10% relative lift needs roughly 50,000 visitors per variant. Wanting to catch a 5% lift instead does not double that — it roughly quadruples it, because the required sample grows with the square of the effect you are chasing. Switch this page to planning mode and it will work out your own number, and how many days that is at your traffic.
Can I stop the test as soon as it hits 95%?
No, and this is the mistake that quietly ruins more experiments than any other. It is called peeking. If you check a running test every day and stop the first time it crosses 95%, you are not accepting a 1-in-20 chance of a false winner — you are taking 20 or 30 separate chances at it, and you will cross that line by luck alone surprisingly often. The fix is to decide the sample size before you start, run to it, and only then look. That is what this page’s planning mode is for. If you have already peeked and stopped early, the result here is real arithmetic on your numbers, but it is optimistic about what those numbers mean.
Should I use a one-sided or two-sided test?
Two-sided, in almost every real case, which is why it is the default here. A two-sided test asks “are these different?”; a one-sided test asks only “is B better?” and cannot tell you that B is worse. One-sided reaches significance on less data, and that is precisely why it gets chosen after the fact by people who want their result to cross the line. The honest use for it is narrow: you have decided in advance that you will ship B or keep A, and a finding that B is actively harmful would change nothing about what you do.
What is “chance B is better”, and why isn’t it just 100% minus the p-value?
Because they answer different questions. The p-value assumes the two variants are identical and asks how surprising your data would be; the probability shown here starts from your data and asks how likely it is that B’s true rate is above A’s. It is a Bayesian figure, computed with a uniform prior — meaning we assume nothing about your conversion rate before seeing your numbers. It is usually the number people thought they were getting when they read a p-value, and it is far easier to say out loud in a meeting: “there is a 98% chance B is better” is a sentence that means what it appears to mean. We show both because the frequentist p-value is still what most teams and most tools report.
Can I compare three or more variants, and is there a sign-up or a limit?
This tool handles exactly two variants, one control and one variation, and says so rather than pretending otherwise. Comparing three or four at once multiplies your chance of a false winner — test four variations against a control and you have four separate opportunities to cross the 95% line by luck — and handling that properly needs a correction this tool does not apply. You can compare each variation against the control here one at a time, as long as you remember that the more comparisons you make, the more sceptical you should be of whichever one wins. As for limits: no account, no email, no paid tier and no cap. Because the calculation happens on your machine rather than ours, there is nothing for us to meter. Nopturnia funds itself through affiliate links on its product buying guides, not through these tools.

Keep it handy

Save it, share it, or tell us what’s missing

Come back to it

Bookmark the page, or pin it to your saved tools so it sits at the top of the tools list on this device.

A number look wrong?

If a figure here disagrees with your testing platform, or you need something this page does not do, tell us. The form opens with the tool already filled in, so you only have to describe what happened.