How much traffic does an A/B test actually need?
A practical way to estimate sample size before you launch an experiment.
The answer is not “run it for two weeks.”
The traffic an A/B test needs depends on four things:
- your current conversion rate
- the smallest change you care about
- your confidence threshold
- your statistical power
If you choose those before the experiment starts, you can estimate the required sample instead of watching the result until it looks convincing.
Start with the baseline
Your baseline is the conversion rate for the control experience.
If 500 of 10,000 visitors convert, the baseline is 5%.
The baseline matters because detecting a small change around a low conversion rate usually takes more observations than detecting a large change.
Decide the minimum detectable effect
The minimum detectable effect, or MDE, is the smallest improvement worth detecting.
Suppose your baseline is 5%.
A 10% relative lift means the variant needs to reach 5.5%.
That is only a 0.5 percentage-point absolute difference:
- Control: 5.0%
- Variant target: 5.5%
- Relative lift: 10%
- Absolute lift: 0.5 percentage points
Small effects require larger samples.
This is why asking for a test to detect “any improvement” is not useful. The smaller the effect you want to detect, the more traffic you need.
Confidence and power solve different problems
A common setup is:
- 95% confidence
- 80% power
Confidence controls how much evidence you require before calling a difference statistically significant.
Power is the probability that the test detects the effect you planned for if that effect is actually there.
Raising either threshold increases the required sample.
Estimate sample size before launch
A sample-size calculation gives you a target for each variation.
For example, a test with:
- 5% baseline conversion
- 10% relative MDE
- 95% confidence
- 80% power
may require tens of thousands of observations per variation.
The exact result depends on the statistical method and assumptions.
Use the Experiment Sample Size Calculator to estimate the requirement for your own baseline and target lift.
Convert sample size into time
Sample size is not the same as duration.
If you need 30,000 visitors per variation and have two variations, you need roughly 60,000 eligible visitors.
At 5,000 eligible visitors per day, that is about 12 days of traffic.
At 500 per day, it is about 120 days.
The Experiment Duration Calculator converts the sample target into a rough calendar estimate.
Do not stop the moment significance appears
Repeatedly checking a test and stopping as soon as the p-value crosses a threshold can increase false positives.
A simple operating rule is to decide the sample target before launch and avoid treating an early crossing as the finish line.
Also account for normal business cycles. A test that technically reaches its sample in three days can still be misleading if weekday and weekend behavior differ.
Statistical significance is not business significance
A statistically significant result can still be too small to matter.
If a change increases conversion from 5.00% to 5.03%, a very large experiment might prove that the difference is real.
That does not mean the change is worth shipping.
Before the test starts, define the effect that would actually change a decision.
A practical planning sequence
Use this order:
- Measure the current baseline.
- Choose the smallest lift worth acting on.
- Pick confidence and power.
- Calculate the sample required per variation.
- Estimate how long that sample will take to collect.
- Run the test without changing the stopping rule.
- Evaluate both statistical and business impact.
The goal is not to make every experiment statistically sophisticated.
It is to decide what evidence you need before you start.