Online Tool Store Online Tool Store
🎯 SEO & Web

· 4 min read

How to Tell if a Conversion Lift Is Real

Heshan Fernando

Co-founder & COO

Heshan Fernando is the Co-founder and Chief Operating Officer of Ceyentra Technologies, where he leads project management, engineering, and research and development strategy. With over nine years of industry experience, he is passionate about transforming complex customer challenges into practical, high-impact solutions. His customer-centric leadership has enabled multidisciplinary teams to consistently deliver secure, scalable, and industry-grade digital products that create lasting business value. View on LinkedIn

Share

How to Tell if a Conversion Lift Is Real

Variant B converts at 3.97% against A’s 3.31%. That is a 20% relative lift, the meeting is on Thursday, and the temptation to call it is considerable.

At those sample sizes the confidence intervals overlap. The difference is suggestive and it is not yet a result.

Relative lift on a small base rate is misleading

A 20% relative improvement sounds substantial. On a 3.31% base rate it is 0.66 percentage points.

Detecting a difference that small reliably needs a large sample, because the noise in a conversion rate at those volumes is of a similar magnitude. That is the whole difficulty with conversion testing: the effects worth having are small in absolute terms, and small effects need a lot of data.

Reporting relative lift is not wrong, and it should always be accompanied by the absolute difference and the interval. “20% lift” and “0.66 points, ±0.5” describe the same result and give very different impressions of how solid it is.

Calculate the sample size before starting

The step that turns testing from a ritual into a method.

Given a base rate and the smallest lift worth detecting, the required sample per variant follows. Run the test to that number and the result means something. Run it until it looks good and it does not.

The relationship is unforgiving: halving the effect you want to detect roughly quadruples the sample needed. Detecting a 20% relative lift on a 3% base rate needs thousands of conversions per variant, not thousands of visitors.

For many sites that means a meaningful test takes weeks. That is a real constraint and the honest response is to test fewer, larger changes rather than many small ones — a redesign of a page is testable where a button colour on a low-traffic site is not.

Base rateDetectSample per variant
3%50% relativeThousands
3%20% relativeTens of thousands
3%5% relativeImpractical for most

Peeking inflates false positives

Checking a running test repeatedly and stopping when it crosses significance is the most common error in the field, and it is worse than it sounds.

Significance testing assumes one look at a predetermined sample size. Looking repeatedly gives many chances to catch a random fluctuation crossing the threshold, and random fluctuations do cross it — the false positive rate rises well above the nominal level, sometimes several times over.

The practical effect is a stream of “winning” variants that do not replicate, and a testing programme whose cumulative effect on conversion is zero while every individual test succeeded.

Two legitimate responses: fix the sample size in advance and look once, or use a sequential testing method designed for continuous monitoring. What does not work is fixed-sample statistics with continuous peeking.

Run for whole weeks

A test stopped mid-week measures a different mix of traffic from the one it started with.

Conversion rates vary systematically by day. Weekday and weekend visitors differ in intent, device and time available, and for many businesses the difference is large.

A test running from Tuesday to the following Monday contains two Mondays and one of everything else, which weights the result toward Monday behaviour. Over a short test that skew can exceed the effect being measured.

Running in whole weeks removes it. It also means the minimum sensible test duration is a week regardless of how quickly the sample accumulates, which is worth knowing when a high-traffic site reaches its sample size in three days.

Common mistakes to avoid

  • Stopping when the result looks favourable.
  • Reporting relative lift without the absolute difference.
  • Running a test with no pre-calculated sample size.
  • Testing during an atypical period — a sale, a holiday, a traffic spike from one source.
  • Running several tests on the same page simultaneously without accounting for interaction.

How to do it with Conversion Rate Calculator

The Conversion Rate Calculator reports intervals alongside rates.

  1. Calculate the required sample before starting the test.
  2. Enter visitors and conversions for each variant.
  3. Read the confidence intervals rather than the point estimates.
  4. If they overlap, keep running — that is not a result yet.

Other analytics tools are in the tools directory.

Frequently asked questions

When can I stop a test?

At the sample size calculated before starting. Stopping when a result looks significant inflates false positives substantially, and it is the most common error in conversion testing.

Why do the intervals overlap when the rates look different?

Because the difference is small relative to the noise at that sample size. A 20% relative lift on a 3% base rate needs tens of thousands of visitors per variant to establish.

Is a 20% lift meaningful?

Only if it is real. A large relative lift on a small sample is usually noise, and relative figures look impressive precisely because the base rate is small.

Final thought

Work out the sample size first and write it down. A test with a stopping rule decided in advance is an experiment; one without is a search for a favourable moment.

Try the free Conversion Rate Calculator

#conversion-rate#ab-testing#statistical-significance#sample-size#online-tools#free-tools