A holdout answers the question correlation cannot: what would have happened without the channel.

Attribution explains how credit is divided, but not what would happen if a channel stopped. We design a controlled holdout, geo split, or spend pulse to answer that causal question.

A pre-registered test design and readout that shows whether the channel is incremental while keeping the uncertainty visible.

A Zeo scientist running a split test on two identical plots, watering only one

Some of the 500+ brands we've worked with

See all references
  • Hyundai
  • Domino’s
  • Shiftdelete
  • Silverline
  • Hepsipay
  • Desa
  • Yolcu360

We lock the design before results appear, preventing the test question or success criteria from shifting afterward. Four stages take the work from an agreed design to a defensible readout with clear boundaries.

How we hold ourselves to it

  • Choose the right test design — We pick a holdout, geo-based, or spend-pulse design based on your channel, volume, and what you can actually control operationally.
  • Define eligibility and success upfront — We specify the eligible population and what counts as the outcome before the test runs, so results can't be reinterpreted after the fact.
  • Watch for contamination — We check for the test group and control group bleeding into each other, a common way incrementality tests quietly produce misleading results.
  • A confidence range, shown plainly — We give you a confidence range and tell you what sample size or duration would be needed for more certainty.
  1. Design the test

    We choose the test structure, define the eligible population and outcome metric, and calculate what sample size or duration you'll actually need.

    Test design document

    AI assist
    Calculates candidate sample sizes for each test duration option.
    Human gate
    Channel owner confirms the design fits real operational constraints.
    Owners
    Measurement Analyst, Channel Owner
    Illustrated figure sketching plans at a drafting table
  2. Set up the holdout or split

    We configure the actual mechanism, a suppressed audience, a geo split, a spend pause, and confirm it's implemented correctly before launch.

    Implementation check

    AI assist
    Checks the holdout configuration against the documented test design.
    Human gate
    Engineering lead confirms the suppression or split is live.
    Owners
    Measurement Analyst, Engineering Lead
    Illustrated figure stacking patterned building blocks
  3. Monitor for contamination

    While the test runs, we watch for cross-contamination between groups and any external factors that could distort the read.

    Monitoring log

    AI assist
    Flags spend or audience overlap between test and control groups.
    Human gate
    Analyst decides whether flagged overlap invalidates the test.
    Owners
    Measurement Analyst
    Illustrated figure watching a monitor full of tracked rows
  4. Report the lift and its range

    We calculate the lift, report the confidence range, and discuss what decision it does and doesn't support.

    Incrementality readout

    AI assist
    Drafts the confidence range from the raw lift calculation.
    Human gate
    Channel owner accepts the range before committing new spend.
    Owners
    Measurement Analyst, Channel Owner
    Illustrated figure presenting a bar chart on an easel

We lock the test design before anyone sees a single result.

Automation sizes the test for each duration option, checks the holdout configuration against the written design, flags spend or audience overlap between groups, and drafts the confidence range from the raw lift. The calls stay human: whether the design fits real operations, whether flagged overlap invalidates the test, and whether the range is good enough to move money.

The result includes the evidence and uncertainty needed to defend the decision it supports.

  • Working document

    Test design document

    The pre-registered design, eligible population, and success metric, written before results existed to bias it.

    Accepted when

    It is dated before the holdout goes live and names the analysis rule in advance.

    Cadence: Dated before the holdout starts

  • Decision memo

    Incrementality readout

    The measured lift with its confidence range, and what decision it supports.

    Accepted when

    The lift is never reported without its confidence range.

    Cadence: At the end of the holdout

  • QA notes

    Contamination and quality notes

    What we checked for cross-contamination and external factors, so the result can be trusted or its caveats understood.

    Accepted when

    Each check states what was examined and what it found, including checks that found nothing.

    Cadence: Checked throughout the run

We call it done when: the lift is reported with its confidence range, contamination has been checked for explicitly, and an inconclusive result is not read as a clean negative.

A real lift test costs volume, time, and a channel you can't touch for a while. A few conditions decide whether that cost is worth paying.

A good fit when

  • Attribution keeps giving the channel credit, but the budget owner still cannot tell whether those conversions would happen without it.
  • Conversion volume and budget flexibility make a real holdout or geo split possible, while the team still lacks causal evidence for the spend decision.
  • A material spend decision is waiting, but the channel owner cannot move the budget until lift and its confidence range are visible.

Better handled as other work when

  • You want a faster, cheaper view of how credit splits across channels, that's Marketing Attribution Modeling.
  • The decision concerns an in-product feature or UX change, so Product Experimentation Measurement should handle it instead of this channel holdout.

If one of these is closer to your situation, start here instead: All Marketing Measurement & Attribution tasks

We call it done when: the eligible population, the success metric, and the test duration are locked and the channel owner confirms the design survives real operations.

  • GeoLift

    designs and analyzes the geo holdout, the causal test this page runs instead of an attribution model

  • Jupyter

    holds the locked test design so it can't be revised once early numbers start looking a certain way

  • Google Analytics

    watched during the live test for contamination bleeding into a holdout market

Tell us which channel you are questioning and what decision depends on the result. We will design a test that can support an evidence-based answer in either direction.
Plan an incrementality test

Attribution divides credit among channels for conversions that have already occurred. Incrementality uses a controlled test to estimate whether the channel caused conversions that would not otherwise have happened. In practice that means running an actual holdout, geo split, or spend pause, which takes real operational coordination to set up. It takes more time and volume because it answers a stronger question.