Creative Testing System
A creative test either has a hypothesis and a decision rule, or it's just two ads running at once.
Swapping images and calling it "testing" doesn't explain why one ad worked. We state the hypothesis and isolate one variable in each cell. The minimum evidence bar is agreed before launch and governs the result. For paid-social accounts already running creative tests, where results aren't consistently written down, compared fairly, or turned into the next test.
A creative test matrix where every win, loss, or inconclusive result leads to a specific next brief and stays available for later rounds.


Some of the 500+ brands we've worked with
See all referencesHow the work runs
The test has to teach the next brief
AI drafts variant directions and flags fatigue. Brand and a compliance reviewer approve every creative. The strategist and creative lead call the winner and decide when the test is done.
How we hold ourselves to it
- Every variant begins with a recorded hypothesis
- One variable changes in each test cell
- The minimum evidence bar is agreed before launch
- AI drafts variant directions. A human approves every creative that runs
Write the hypothesis before the variant
State what you believe and why (a specific hook, format, or proof point) before any creative gets built, so the test has something real to confirm or reject.
Written hypothesis naming the hook, format or proof point under test and the reason behind it.
- AI assist
- Clusters past results and research into a draft hypothesis list ranked by likely impact.
- Human gate
- Creative lead and strategist approve which hypotheses get built.
- Owners
- Creative lead + strategist


Build a cell that isolates the variable
Construct the test so the one thing you're testing is the only thing that changed. Placement, audience, and budget logic all stay the same across cells.
Test cell where placement, audience and budget logic are held constant across variants.
- AI assist
- Checks the proposed test cell against the isolation requirement and flags confounded setups.
- Human gate
- Strategist signs off on the cell design before build.
- Owners
- Strategist + media specialist


Set the evidence bar before launch
Agree the sample size, duration, and confidence level that will count as a real result before anyone sees early numbers and gets tempted to call it early.
Agreed sample size, duration and confidence level recorded before any numbers arrive.
- AI assist
- Calculates a rough sample-size estimate from historical volume and flags an unrealistic timeline.
- Human gate
- Strategist and client approve the evidence bar.
- Owners
- Strategist + client


Produce and validate the creative
Build the approved variants within brand and platform policy, and check rights and claims before anything launches.
Approved variants cleared against brand rules, platform policy, rights and claims.
- AI assist
- Drafts reviewable variant copy and flags policy or claim risks against the brief.
- Human gate
- Brand and a compliance reviewer approve every variant before launch.
- Owners
- Brand owner + compliance reviewer


Run it, and keep watching for fatigue
Let the test run to its agreed threshold, and separately track frequency and engagement decay so a still-running test doesn't quietly fatigue mid-flight.
Run record with frequency and engagement decay tracked separately from the test result.
- AI assist
- Flags rising frequency or falling engagement against the fatigue threshold.
- Human gate
- Creative lead decides whether to refresh, pause, or let a fatiguing variant finish.
- Owners
- Creative lead + media specialist


Call the result and write the next brief
Judge the test against the evidence bar set in step three (win, loss, or inconclusive), and turn whatever was learned into a specific hypothesis for the next round.
Verdict against the evidence bar plus the next hypothesis it produced.
- AI assist
- Drafts the result summary and a candidate next-brief from the test data.
- Human gate
- Strategist and creative lead approve the call and the next brief.
- Owners
- Strategist + creative lead


What lands with your team
A matrix that remembers what the account learned
The team doesn't have to remember whether a hook was tested months ago. The matrix keeps the result and the brief that followed it.


Creative test matrix
Every test's hypothesis, cell design, evidence bar, and result in one place, so nothing gets retested by accident.
Accepted when
Every test's hypothesis, cell design, evidence bar and result sit in one place, so a variant is never retested by accident.
Cadence: Per creative round


Fatigue tracking log
Frequency and engagement trend per active creative, with the threshold that triggers a refresh.
Accepted when
Each active creative carries its frequency and engagement trend and the threshold that triggers a refresh.
Cadence: Reviewed weekly


Result and confidence record
What each test actually proved, at what confidence, and what it didn't answer.
Accepted when
Each record states what the test proved, at what confidence, and what it did not answer.
Cadence: Per test close


Next-brief queue
A running list of hypotheses ready to build, ranked by what the last round of evidence supports.
Accepted when
The queue is ranked by what the last round of evidence supports, not by which idea arrived most recently.
Cadence: After every test call
Before the work starts
When this work is the right next step
Two ads ran. One had more clicks. That's not a test result. Most "creative testing" is really just running a few ad variations and picking whichever one reports the lowest cost after a few days. Without a stated hypothesis, an isolated variable, and an evidence bar set in advance, that's a coin flip wearing a spreadsheet, and it teaches the account nothing it can reuse next month.
A good fit when
- Creative refreshes happen regularly, but hypotheses, cell designs, and results stay scattered across the account, so the next round starts without a usable record.
- Each test cell gets enough spend to reach the evidence bar set before launch, yet early cost swings still tempt the team to call a winner.
- A test ends with a result, but no approved next brief follows, so the account keeps inventing fresh variants instead of using what it learned.
- Multiple creative variants run in the same ad set with no stated hypothesis behind any of them.
- A test gets called "won" after a day or two, before enough evidence has actually accumulated.
- The same creative keeps running past the point where frequency and falling engagement say it's fatigued.
- A test result never turns into a specific next brief: it just ends, and the next batch starts from scratch.


Better handled as other work when
- The budget cannot give two isolated variants a meaningful sample, so the platform's first cost difference would be mistaken for evidence.
- Creative changes are reactive and carry no written hypothesis, so a lower cost gives the next brief nothing it can explain or reuse.
- No one will own the result or write the next brief, so the next creative round starts from scratch.
People who own a single channel
Paid Search, Paid Social, CRO, and Programmatic each run under a named owner at Zeo. The consultants below are matched to the channel this page is about, so you can see who you'd actually work with.

Serap Yurtvermez
Performance Marketing Team Lead

Abdullah Tanıdır
Performance Marketing Team Lead

Sevda Yurtvermez
Performance Marketing Team Lead

İlker Emir
Senior Performance Marketing Executive

Onur Durdağı
Performance Marketing Executive

Ozan Ketenci
VP of Consulting & Strategy

Burak Pehlivan
Co-founder & CEO
Tools we use
Tools behind this work
Smartly.ioproduces the isolated variant across every platform in the test from one build, not a re-export per network
Madgicxwatches for fatigue after the result is called, not just during the test itself
Bring your current tests
Stop guessing which ad won and why





















