How much traffic you need before an A/B test means anything
Detecting a 10% lift on a 3% conversion rate takes 51,700 visitors per variant. Run that math before launch, so you test something you can measure.
At a 3% conversion rate, catching a 10% lift needs roughly 51,700 visitors per variant. Two arms, so 103,000 visitors before the test can tell you anything. If the page you want to test gets 8,000 visitors a month, the answer to how much traffic you need for an A/B test is: more than you have, and you should test a bigger change instead.
How much traffic you need for an A/B test
Four inputs decide the number. Your baseline conversion rate, the smallest lift worth catching (the minimum detectable effect, or MDE), your confidence level, and your power. Hold confidence at 95% and power at 80%, which is what almost everyone uses, and the rest collapses into one line of arithmetic: visitors per variant equals 16 x p x (1 - p) divided by the absolute lift squared, where p is your baseline rate.
Work the example. Baseline 3%, target lift 10% relative, so the absolute lift is 0.3 percentage points. That gives 16 x 0.03 x 0.97 divided by 0.003 squared, or about 51,700 visitors per variant. Any sample size calculator returns the same figure with a nicer interface. The point of doing it by hand once is that you stop being surprised by the answer.
Small wins cost a lot to measure
Keep the 3% baseline and change only the lift you want to detect:
- 20% relative lift: about 12,900 visitors per variant
- 10% relative lift: about 51,700
- 5% relative lift: about 207,000
Halve the effect and the sample roughly quadruples. That one relationship should shape your whole test roadmap. Button colour and headline tweaks move conversion by a few percent at best, which puts them out of reach for most B2B sites forever. Rewriting the offer, cutting three form fields, changing what the page asks for: those move numbers far enough that a normal traffic budget can see them.
The baseline hurts too. At a 1% conversion rate, that same 10% lift needs about 158,400 visitors per variant. Low-converting pages are the hardest places to learn anything, which is worth knowing before you pick one as your first test.
Lock the sample size before you launch
Write the number down, then stop looking until you hit it. Checking results as they come in and stopping when the line goes green is the fastest way to ship a win that was never there. Evan Miller's walkthrough of repeated significance testing puts a figure on it: peek continuously and your real false positive rate lands around 26.1%, five times the 5% you thought you were running. Peek ten times and what reads as 1% significance is really 5%.
CXL ran the demonstration that makes it concrete. A thousand A/A tests, identical pages on both sides, nothing to find. 771 of them crossed 90% significance at some point during the run, and 531 crossed 95%. Every one of those would have been called a winner by someone watching the dashboard.
Deciding the stopping rule in advance removes the problem. If you want to look early, the honest versions are sequential testing or a Bayesian setup, both of which price the peeking into the math.
Run it in whole weeks
Sample size tells you how many. Duration tells you who. Tuesday traffic and Saturday traffic are different people with different intent, and a test that ends mid-week has quietly overweighted one of them. Two full weeks is the usual floor, and if your buying cycle runs monthly, use two of those instead.
Payday weeks, end-of-quarter pushes, a competitor's launch, a public holiday: each one can move conversion more than your variant does. Whole cycles average them out across both arms.
What to do when the traffic isn't there
Most B2B and SaaS sites don't have testing traffic on their money pages, and that's fine. Four moves that work:
Test bigger changes. If you can only detect a 20% lift, only run tests that could plausibly produce one. A new offer, a different page structure, a shorter form.
Move the metric up the funnel. Demo requests might be 40 a month while clicks into the pricing table are 4,000. Testing the earlier step gets you an answer in weeks, as long as you check afterward that the later step moved too.
Pool the surface. One test running across every pricing page or every blog CTA gathers the sample several times faster than one page on its own.
Buy the sample. Paid traffic to a test page turns a twelve-month experiment into a three-week one, and you already know what that traffic costs you. Budget it as research, not acquisition.
When none of those apply, decide with judgement and qualitative evidence. Session recordings, five customer interviews, a support inbox read end to end. That's weaker evidence than a clean experiment and much stronger evidence than an underpowered one. The related trap is measuring the wrong thing well, which we covered in the growth metric most SaaS teams track wrong.
The test you can't power is a test you shouldn't run. Do the arithmetic first and you spend the next quarter on questions your traffic can answer.
If you want help building a test roadmap your traffic can support, that's the kind of work we do.