THE CREATIVE TESTING FRAMEWORK
Test everything, trust nothing. Creative testing is how I settle arguments about creative: let the audience vote with their clicks, not the loudest voice in the room. As David Ogilvy put it: “If it doesn’t sell, it isn’t creative.”
Not to Be Confused With
🎨 Brand guidelines
Guidelines are the constraints; testing is what happens inside them. Netflix’s Kelly Bennett turned off all direct response marketing to reset a culture where sexy images and price discounts were harming the long-term brand. Creative must have constraints. Then you test freely within them.
📊 Testing vs optimisation
Testing discovers what works; optimisation scales it. A winning test tells you this creative beat that creative. It does not tell you why, or whether a bolder idea would have beaten both.
What to Test
Copy & CTAs
Headlines, body text, button text. Small words, big impact. I once found a 40% lift hiding in a translation: the French version of our page was converting 40% above the global average, and the difference came down to “Advertise” versus “Create an Ad”.
“Advertise” → “Create an Ad”: 40% lift
Images & Video
Product shots vs lifestyle, static vs motion, with or without people.
Personalization
Name in subject line, contextual logos, behavioral targeting.
Ad Format
Single image, carousel, video, stories format. Each reaches differently.
Placement & Prominence
Feed vs sidebar, above vs below fold, mobile vs desktop. Moving a link from the bottom of the page to the top gave us 30%.
Link bottom → top: 30% lift
Social Proof
Friend endorsements, ratings, testimonials, user counts.
Persistence & Frequency
How many times to show before sequencing to new creative (impression discounting).
Landing Page
What happens after the click matters as much as the click itself.
How to Test
-
Isolate each test. Multivariate tests are fine, but don’t cross-contaminate experiments. Keep each test cleanly separated so you know what caused the difference.
-
Pre-register test length and sustain significance. Decide the test duration upfront. One day of statistical significance is not enough: you need sustained significance over the pre-registered period. Don’t peek and kill early.
-
Business significance, not just statistical. A statistically significant 0.01% lift doesn’t matter if nobody notices.
-
Run long enough. Account for day-of-week effects, time zones, seasonal patterns.
-
Document everything. Build institutional knowledge. Today’s loser insight is tomorrow’s winning hypothesis.
Reading Results
✅ This Result Matters
- Double-digit percentage change
- Visible impact on business metrics
- Finance would care if it went away
- Replicable across regions/audiences
This Result Doesn’t Matter
- Needs a microscope to detect
- Statistically significant but tiny
- Only affects a niche segment
- Can’t be explained simply
Impression Discounting
The Results > Awards Rule
Results Over Awards
To be clear, I am not against awards. My teams won 88 of them in 2024, including a Cannes Grand Prix and five other Grand Prix, and we were named best in-house agency at the London International Awards. Awards can recognise genuinely great work.
“My number one sign to look for in a creative agency is that they focus on their results over their awards. Most effective digital direct response campaigns don’t even submit for industry awards. Marketing awards don’t tend to correlate closely with the most effective work.”
Know Its Limits
My position: testing refines ideas, it does not have them. No amount of A/B testing produces the breakthrough concept; tests refine, humans leap. If you test everything but create nothing bold, you will converge on competent mediocrity.
And measurement can become a tyranny. I test to learn, not to avoid blame. A team that only ships what tests well stops taking the creative risks that produce the biggest wins: the ones no test would have predicted.
