All notes

Busting Myths

CRO is not just A/B testing (and the red button will not save you)

The short version: conversion rate optimization is not A/B testing. Testing is the last, smallest step in a research-led process, and most published tests never reach statistical significance. The real levers are value proposition clarity, trust, and the offer itself. A red button will not fix a page that gives visitors no clear reason to say yes.

Ask most marketers what conversion rate optimization means, and the answer arrives fast: run a test, pick a winner, repeat. It is a tidy story, and mostly wrong. Two myths travel together here. The first is that CRO equals testing. The second is that the tests worth running are cosmetic: a button color, a headline tweak, a CTA word swap. Neither survives the data.

What is conversion rate optimization, if it is not A/B testing?

CRO is the discipline of finding and removing the real reasons a visitor does not take the action a page wants, then confirming the fix with evidence. Testing sits at the end of that chain, as verification. Most of the work happens earlier: reading analytics funnels for where people leave, watching session recordings for where they hesitate, and building a specific, falsifiable hypothesis before any test goes live. Skip that work and a test is not an experiment. It is a guess with a sample size.

How many A/B tests actually produce a winner?

Fewer than most teams assume, and the gap between the myth and the number is exactly why testing alone cannot carry a CRO program.

SourceTests examinedStatistically significant win rate
Optimizely127,000+ experiments12 percent on the primary metric (Optimizely)
ConversionTeam2,288 audited tests19.1 percent per test (ConversionTeam, to re-verify)

Independent studies land anywhere from roughly 10 to 22 percent, depending on how strictly "win" is defined, but they agree on the shape: most tests are flat, not victories. That is not an argument against testing. It is an argument against spending scarce testing capacity on low-confidence ideas instead of the hypotheses research already points to.

What actually moves the number, if not the test itself?

Three things, in this order.

  1. Value proposition clarity. Whether a visitor understands, in seconds, what the offer is, who it is for, and why it beats the alternative.
  2. Trust. Whether the visitor believes the page and the business behind it before being asked to commit anything.
  3. The offer. The actual terms on the table: price, risk, packaging, and what is being asked of the visitor relative to what they have earned.

None of these are cosmetic, and none are fixed by swapping a hex code. NextAfter's published experiment log includes a donation-page test where a sharper value proposition lifted conversion by roughly 41 percent, on a sample honest enough to report the result narrowly missed the standard 95 percent confidence bar (NextAfter, to re-verify): exactly the discipline real optimization work requires. On trust, Baymard Institute found that 19 percent of shoppers had abandoned a checkout specifically over payment trust (Baymard Institute). Neither lever shows up in a "before" and "after" button screenshot.

Why did the red button become gospel?

Because one case study was easy to repeat and nobody checked the context. In 2011, HubSpot ran a red button against a green one on a sample of a little over 2,000 visits, and red won by 21 percent (HubSpot, to re-verify, original small-sample test). Green was HubSpot's own dominant brand color, used everywhere on that page, so visitors had gone blind to it. Red was the only warm color on the screen. The lever was contrast against a color-saturated page, not the hue itself, a point conversion researchers including CXL have made consistently since (CXL, to re-verify). A result that specific, from a page that specific, was never a universal law. It just travelled well.

Where does testing actually belong in the process?

Last, and only once research has produced a hypothesis worth the traffic it will cost to test. One CRO agency puts a number on the split: practitioners worth hiring spend roughly 70 percent of their time analyzing data and 30 percent testing, a single-source estimate rather than a survey finding, but directionally consistent with everything above (Strategyc, to re-verify). It also explains why testing gets the credit: a test produces a clean number, research produces judgment, and judgment is harder to put in a slide. A widely repeated, Econsultancy-sourced figure names a longer-standing version of the same imbalance: for every 92 dollars spent acquiring a visitor, only 1 goes toward converting them. Its original report is hard to trace at this remove, so treat it as industry lore, not a fresh study (to re-verify). True or approximate, the direction holds: acquisition gets the budget, and conversion gets what testing can squeeze from the rest.

Value proposition, trust, and offer are where that budget should go first. Testing confirms the fix. It does not find it.

Frequently asked questions

No. Testing is one verification step inside CRO, which starts with research into why visitors do not convert. Most of the gains happen before a test ever goes live.

Roughly 12 to 19 percent, by two large-scale analyses across 127,000-plus and 2,288 audited tests. Independent studies land in a 10 to 22 percent range. Most tests are flat or inconclusive, not winners.

Only through contrast, not the hue itself. The famous red-versus-green test behind this myth used a page where green was already the dominant brand color, so red simply stood out. Contrast is worth testing. Color for its own sake is not.

It produced a clean, dramatic number from one page, and that number travelled further than its context did. The lift came from contrast on a color-saturated page, a nuance dropped every time the result got repeated as a universal rule.

A specific, falsifiable hypothesis built from evidence: analytics showing where visitors leave, session recordings showing where they hesitate, and research into what confuses them. A test without that is a guess wearing a sample size.

Value proposition clarity, trust, and the offer itself: whether a visitor understands the offer and why it beats the alternative, whether they believe the site, and whether the terms match what they have earned.

There is no universal ratio, but practitioners who take it seriously weight it heavily toward research and diagnosis, with testing as the final step that confirms a hypothesis rather than generates one.

A long-circulated figure claims a 92-to-1 ratio. Its original source is hard to trace, so treat it as industry lore, not a fresh study. The direction it points to, acquisition getting the budget and conversion getting the leftovers, matches practitioner experience.

The answer to what this is, who it is for, and why it beats the alternative, understood within seconds. A published experiment recorded a roughly 41 percent lift from sharpening one, a scale no button hue has ever matched.

Long enough to build a hypothesis from evidence, not a hunch, with a clear reason to expect the change will move the metric.

No. A redesign changes many variables at once, so when the number moves, nobody can say which change did it.

Treating the test as the strategy, not the last step of one. Fix clarity, trust, and the offer first, then test the change with the best odds.