CRO is niet alleen A/B-testen (en de rode knop redt u niet)
The short version: conversion rate optimization is not A/B testing. Testing is the last, smallest step in a research-led process, and most published tests never reach statistical significance. The real levers are value proposition clarity, trust, and the offer itself. A red button will not fix a page that gives visitors no clear reason to say yes.
Ask most marketers what conversion rate optimization means, and the answer arrives fast: run a test, pick a winner, repeat. It is a tidy story, and mostly wrong. Two myths travel together here. The first is that CRO equals testing. The second is that the tests worth running are cosmetic: a button color, a headline tweak, a CTA word swap. Neither survives the data.
What is conversion rate optimization, if it is not A/B testing?
CRO is the discipline of finding and removing the real reasons a visitor does not take the action a page wants, then confirming the fix with evidence. Testing sits at the end of that chain, as verification. Most of the work happens earlier: reading analytics funnels for where people leave, watching session recordings for where they hesitate, and building a specific, falsifiable hypothesis before any test goes live. Skip that work and a test is not an experiment. It is a guess with a sample size.
How many A/B tests actually produce a winner?
Fewer than most teams assume, and the gap between the myth and the number is exactly why testing alone cannot carry a CRO program.
| Source | Tests examined | Statistically significant win rate |
|---|---|---|
| Optimizely | 127,000+ experiments | 12 percent on the primary metric (Optimizely) |
| ConversionTeam | 2,288 audited tests | 19.1 percent per test (ConversionTeam, to re-verify) |
Independent studies land anywhere from roughly 10 to 22 percent, depending on how strictly "win" is defined, but they agree on the shape: most tests are flat, not victories. That is not an argument against testing. It is an argument against spending scarce testing capacity on low-confidence ideas instead of the hypotheses research already points to.
What actually moves the number, if not the test itself?
Three things, in this order.
- Value proposition clarity. Whether a visitor understands, in seconds, what the offer is, who it is for, and why it beats the alternative.
- Trust. Whether the visitor believes the page and the business behind it before being asked to commit anything.
- The offer. The actual terms on the table: price, risk, packaging, and what is being asked of the visitor relative to what they have earned.
None of these are cosmetic, and none are fixed by swapping a hex code. NextAfter's published experiment log includes a donation-page test where a sharper value proposition lifted conversion by roughly 41 percent, on a sample honest enough to report the result narrowly missed the standard 95 percent confidence bar (NextAfter, to re-verify): exactly the discipline real optimization work requires. On trust, Baymard Institute found that 19 percent of shoppers had abandoned a checkout specifically over payment trust (Baymard Institute). Neither lever shows up in a "before" and "after" button screenshot.
Why did the red button become gospel?
Because one case study was easy to repeat and nobody checked the context. In 2011, HubSpot ran a red button against a green one on a sample of a little over 2,000 visits, and red won by 21 percent (HubSpot, to re-verify, original small-sample test). Green was HubSpot's own dominant brand color, used everywhere on that page, so visitors had gone blind to it. Red was the only warm color on the screen. The lever was contrast against a color-saturated page, not the hue itself, a point conversion researchers including CXL have made consistently since (CXL, to re-verify). A result that specific, from a page that specific, was never a universal law. It just travelled well.
Where does testing actually belong in the process?
Last, and only once research has produced a hypothesis worth the traffic it will cost to test. One CRO agency puts a number on the split: practitioners worth hiring spend roughly 70 percent of their time analyzing data and 30 percent testing, a single-source estimate rather than a survey finding, but directionally consistent with everything above (Strategyc, to re-verify). It also explains why testing gets the credit: a test produces a clean number, research produces judgment, and judgment is harder to put in a slide. A widely repeated, Econsultancy-sourced figure names a longer-standing version of the same imbalance: for every 92 dollars spent acquiring a visitor, only 1 goes toward converting them. Its original report is hard to trace at this remove, so treat it as industry lore, not a fresh study (to re-verify). True or approximate, the direction holds: acquisition gets the budget, and conversion gets what testing can squeeze from the rest.
Value proposition, trust, and offer are where that budget should go first. Testing confirms the fix. It does not find it.
Veelgestelde vragen
Nee. Testen is één verificatiestap binnen CRO, dat begint met onderzoek naar waarom bezoekers niet converteren. Het grootste deel van de winst wordt behaald voordat een test ooit live gaat.
Ongeveer 12 tot 19 procent, volgens twee grootschalige analyses op meer dan 127.000 respectievelijk 2.288 geauditeerde tests. Onafhankelijke studies komen uit op een bereik van 10 tot 22 procent. De meeste tests laten geen effect zien of geven geen uitsluitsel, het zijn geen winnaars.
Alleen via contrast, niet via de kleurtint zelf. De beroemde rood-tegen-groen-test achter deze mythe gebruikte een pagina waarop groen al de dominante merkkleur was, waardoor rood simpelweg opviel. Contrast is het testen waard. Kleur op zichzelf niet.
De test leverde een helder, spectaculair cijfer op van één pagina, en dat cijfer verspreidde zich verder dan de context waarin het ontstond. De stijging kwam door contrast op een kleurverzadigde pagina, een nuance die steeds verloren ging wanneer het resultaat werd herhaald als universele regel.
Een specifieke, falsifieerbare hypothese, opgebouwd uit bewijs: analytics die laten zien waar bezoekers afhaken, sessieopnames die laten zien waar ze aarzelen, en onderzoek naar wat hen in verwarring brengt. Een test zonder dat is een gok die zich voordoet als steekproef.
Een heldere waardepropositie, vertrouwen en het aanbod zelf: of de bezoeker het aanbod begrijpt en waarom het beter is dan het alternatief, of die de site vertrouwt, en of de voorwaarden overeenkomen met wat die heeft verdiend.
Er bestaat geen universele verhouding, maar wie het serieus aanpakt, legt het zwaartepunt bij onderzoek en diagnose, met testen als laatste stap die een hypothese bevestigt in plaats van genereert.
Een veelgebruikt cijfer noemt een verhouding van 92:1. De oorspronkelijke bron is nauwelijks te achterhalen, dus behandel het als branchefolklore, niet als een recente studie. De richting die het aangeeft, acquisitie krijgt het budget en conversie krijgt de kruimels, komt overeen met wat vakmensen in de praktijk ervaren.
Het antwoord op wat dit is, voor wie het is, en waarom het beter is dan het alternatief, begrepen binnen enkele seconden. Een gepubliceerd experiment registreerde een stijging van ongeveer 41 procent na het aanscherpen van een waardepropositie, een schaal die geen enkele knopkleur ooit heeft geëvenaard.
Lang genoeg om een hypothese op te bouwen uit bewijs, niet uit een onderbuikgevoel, met een duidelijke reden om te verwachten dat de verandering de metriek zal beïnvloeden.
Nee. Een redesign verandert veel variabelen tegelijk, dus wanneer het cijfer verandert, kan niemand zeggen welke wijziging daarvoor zorgde.
De test behandelen als de strategie, in plaats van als de laatste stap ervan. Verbeter eerst de helderheid, het vertrouwen en het aanbod, en test daarna de wijziging met de beste kans van slagen.