A/B Test Design & Analysis
A test result is only as credible as the design decisions made before anyone sees the data.
Before anyone sees a variant, we agree on the sample size, minimum detectable effect, randomization unit and decision rule. The test then runs to that registered point. Refreshing the dashboard every day doesn't replace a stopping rule, and calling the first green number a win raises the risk of a false positive. You have one experiment to run and need the statistical design settled before the build starts, wherever the test sits on the site.
A test result that can be interpreted against decisions documented before exposure, with any later deviations visible in the record.


Some of the 500+ brands we've worked with
See all referencesHow the work runs
Set the design, run the test, read the result
Most of the important calls happen before launch. AI can help with calculations and a first draft of the readout from approved inputs, but the CRO strategist interprets the result. AI never declares a winner.
How we hold ourselves to it
- Set sample size and minimum detectable effect before exposure begins
- The randomization unit matches how people actually experience the page
- The pre-registered decision rule determines the stopping point
- Report the confidence interval with the result
Define the minimum effect worth detecting
Set the smallest change in the primary metric that would matter to the business. A smaller detectable number can still be too slight to justify a testing slot.
Agreed minimum detectable effect, written down with the business reason it matters.
- AI assist
- Drafts a minimum-detectable-effect worksheet from historical baseline data.
- Human gate
- The CRO strategist and client owner agree on the threshold before sample size is calculated.
- Owners
- CRO strategist + client owner


Calculate sample size and duration before launch
Use the agreed minimum effect, baseline conversion rate, and desired significance and power to calculate the required sample size, then translate that into an expected run duration given current traffic.
Required sample size and expected run duration, calculated from the agreed effect, baseline rate, significance and power.
- AI assist
- Runs the sample-size calculation and flags it if the resulting duration looks impractical.
- Human gate
- The experiment owner signs off on the calculated duration before build starts.
- Owners
- Experiment owner + analytics owner


Choose a randomization unit that matches real behavior
Decide whether visitors, sessions, or logged-in users are the unit of assignment, based on how people actually move through the page. Get this wrong and the same person can land in both variants.
Recorded randomization unit and the behavioral reason it was chosen.
- AI assist
- Flags a mismatch between the proposed randomization unit and known user behavior patterns.
- Human gate
- The CRO strategist approves the randomization unit before configuration begins.
- Owners
- CRO strategist + engineering


Pre-register the decision rule
Before launch, write down exactly which result, metric, and threshold will trigger a ship, iterate, or stop decision. This prevents anyone from choosing the rule after seeing which one favors the outcome they wanted.
Signed decision rule naming the metric, threshold and direction that trigger ship, iterate or stop.
- AI assist
- Drafts the decision rule document from the agreed metric and threshold.
- Human gate
- The CRO strategist and client owner sign the decision rule before exposure begins.
- Owners
- CRO strategist + client owner


Run to the pre-registered point
Let the test reach its calculated sample size or duration before analysis. A favorable dashboard reading does not change the stopping point. Early stopping inflates the false-positive rate.
Run log showing the test reached its planned sample size or duration before anyone read it.
- AI assist
- Tracks progress against the pre-registered sample size and flags any early-stop request for review.
- Human gate
- The experiment owner approves any deviation from the pre-registered stopping point, with the reason recorded.
- Owners
- Experiment owner + analytics owner


Analyze against the one pre-registered rule
Report the result with its confidence interval against the decision rule written before the test started. Movement in a secondary metric is labeled separately from the pre-registered call.
Readout stating the result and its confidence interval against the pre-registered rule, with secondary movement labeled separately.
- AI assist
- Drafts the analysis readout against the pre-registered rule, labeling any secondary-metric movement separately.
- Human gate
- The CRO strategist approves the readout. The client owns the shipping decision.
- Owners
- CRO strategist + client owner


What lands with your team
What we keep with the result
The numbers make sense only when the original design is still beside them. These records show what the team agreed before launch and what changed during the run.


Design worksheet
Minimum detectable effect, baseline rate, sample size, and expected duration, calculated before the test launches.
Accepted when
The minimum detectable effect, baseline rate, sample size and duration are all filled in and dated before exposure begins.
Cadence: Before every test


Randomization and assignment record
The chosen randomization unit and the reasoning behind it, plus a running check that assignment held steady across sessions.
Accepted when
The unit of assignment is named with its reasoning, and the sample-ratio check shows assignment held for the full run.
Cadence: Set before launch


Pre-registered decision rule
The exact metric, threshold, and direction that will decide ship, iterate, or stop, written and signed before exposure begins.
Accepted when
The metric, threshold and direction are written and signed before the first visitor is exposed.
Cadence: Locked before launch


Analysis readout with confidence interval
The result reported against the pre-registered rule, with a confidence interval attached and any secondary-metric movement labeled separately.
Accepted when
The headline call cites the pre-registered rule, carries a confidence interval, and keeps secondary-metric movement in a separate section.
Cadence: At sample size
Before the work starts
When this work is the right next step
Many A/B tests lose their validity before the data is even available. A test can run for weeks and still tell you nothing. The idea may be sound, but the design fails if nobody sets a sample size before starting, the dashboard is checked daily until a green number appears, or the randomization unit does not match how people browse. The math doesn't forgive shortcuts taken during design.
A good fit when
- Your page has enough traffic to reach the calculated sample size, so the test can close within a window the change is worth waiting for.
- You need the sample size, run duration, and stopping rule settled before launch, because changing them mid-test would break the registered design.
- The shipping decision can stay open until the pre-registered point, so an early dashboard swing does not decide which variant goes live.
- The test gets called the moment the dashboard shows a "significant" green number, whenever that happens to land.
- Nobody calculated a minimum detectable effect or sample size before the test launched.
- Visitors get re-randomized on every page view instead of staying in one variant for the session or visit.
- Five metrics get checked and whichever one turned green becomes "the result."


Better handled as other work when
- Traffic is too thin to reach the calculated sample size before the value of the change has passed, so research or a controlled release fits better.
- Daily dashboard checks are expected to decide the winner, but this test holds the read until the registered stopping rule is met.
- Several page changes move together, but no interaction plan shows which one caused the variant result.
People who own a single channel
Paid Search, Paid Social, CRO, and Programmatic each run under a named owner at Zeo. The consultants below are matched to the channel this page is about, so you can see who you'd actually work with.

Serap Yurtvermez
Performance Marketing Team Lead

Abdullah Tanıdır
Performance Marketing Team Lead

Sevda Yurtvermez
Performance Marketing Team Lead

İpek Ezer
Performance Marketing Executive

Onur Durdağı
Performance Marketing Executive

Zafer Yıldız
Web Analytics Manager

Metehan Urhan
New Business & Partnership Manager
Tools we use
Tools behind this work
VWOruns the test on a sequential model, so an early look doesn't invalidate the read
Google Analyticschecks how the account's own data is scoped before the randomization unit gets locked in
Settle the design before the build
Find out whether this test is feasible



















