The rules that decide the test are written down before the first visitor is assigned.

A test loses credibility when people check it every day and stop as soon as the result looks favorable. We set the decision rules before launch and apply them consistently, whatever the result shows.

A defensible test design and readout, with guardrails that reveal when one metric improves at another's expense.

Two Zeo testers at an A/B door pair, measuring foot traffic

Some of the 500+ brands we've worked with

See all references
  • Domino’s
  • Silverline
  • Odeabank
  • Duru
  • Gusto
  • Elele

We agree on the rules before anyone can adjust them to fit the emerging data. Four steps take the experiment from a defined hypothesis to a defensible decision.

How we hold ourselves to it

  • What counts as success, decided before anyone looks — We define the hypothesis and primary success metric before anyone sees the results.
  • Enough traffic that a real effect won't hide in the noise — We estimate the traffic or duration needed to detect a meaningful effect, reducing the temptation to stop early because of noise.
  • A metric that must not get worse — We choose guardrail metrics upfront so an improvement in the primary metric cannot conceal harm elsewhere.
  • When and how the test gets read, fixed in advance — We agree when and how to analyze the test, removing the option to stop at a convenient-looking result.
  1. Design the experiment

    We define the hypothesis, unit of assignment, sample size, guardrails, and stopping rule before launch.

    Experiment design document

    AI assist
    Drafts a sample-size estimate from your traffic data.
    Human gate
    Your product owner confirms the primary metric.
    Owners
    Experiment Analyst, Product Owner
    Illustrated figure sketching plans at a drafting table
  2. Verify the setup

    Before the full launch, we verify that assignment works as intended and that the test has not already been contaminated.

    Setup verification

    AI assist
    Checks the assignment split against the intended ratio.
    Human gate
    Analyst confirms the test isn't already contaminated.
    Owners
    Experiment Analyst, Engineering Lead
    Illustrated figure reading an oversized measurement dial
  3. Monitor without peeking-driven decisions

    During the run, we monitor technical health without using interim performance to make stop or continue decisions.

    Monitoring log

    AI assist
    Flags technical anomalies during the run automatically.
    Human gate
    No one calls the test before the rule date.
    Owners
    Experiment Analyst, Product Owner
    Illustrated figure watching a monitor full of tracked rows
  4. Call the result

    We analyze the outcome using the agreed rule and report it plainly, including a null result when that is what the test produced.

    Experiment readout

    AI assist
    Drafts the readout from the pre-agreed analysis rule.
    Human gate
    Analyst confirms every guardrail before recommending a rollout.
    Owners
    Experiment Analyst, Product Owner
    Illustrated figure presenting a bar chart on an easel

Sample size, guardrails, and stopping rules are fixed before launch.

Automation estimates the sample size from your traffic, checks the assignment split against the intended ratio, flags technical anomalies during the run, and drafts the readout from the pre-agreed rule. The discipline is human: nobody calls the test before the rule date, and an analyst clears every guardrail before a rollout is recommended.

The readout remains useful whether the result is positive, negative, or inconclusive.

  • Working document

    Experiment design document

    The hypothesis, sample size calculation, guardrails, and stopping rule, written before results existed.

    Accepted when

    Its timestamp precedes the first assignment, so the rules cannot have been chosen after the fact.

    Cadence: Written before assignment starts

  • Decision memo

    Experiment readout

    The result, analyzed according to the pre-agreed rule, with guardrail performance included.

    Accepted when

    The primary metric and every guardrail are reported together, whichever way the result went.

    Cadence: At sample size

  • Reference document

    Decision record

    What was decided based on the result and why, so it's traceable later.

    Accepted when

    The record states what was decided and which part of the readout decided it.

    Cadence: One per experiment

We call it done when: the readout follows the pre-registered rule, every guardrail is reported alongside the primary metric, and a decision to ship nothing is written up as carefully as a win.

This work fits teams that need a controlled product experiment and an honest readout.

A good fit when

  • You're testing a specific in-product feature or change and need a properly designed experiment rather than an informal split.
  • Past tests have been stopped early or re-analyzed after the fact, and results no longer feel trustworthy.
  • You need guardrail metrics to catch a change that helps one number while quietly hurting another.

Better handled as other work when

  • You're testing marketing spend or channel effectiveness rather than an in-product change. That is Incrementality & Lift Measurement.
  • You need us to build and ship the treatment itself. We design the test and read the result. Your product or engineering team implements the change.

If one of these is closer to your situation, start here instead: All Product & Customer Analytics tasks

We call it done when: the hypothesis, primary metric, sample size, guardrails, and stopping rule are recorded and the product owner has signed off on the primary metric.

  • Evan Miller's A/B Tools

    the neutral calculator and reference essay behind the fixed sample size and stopping rule

  • Amplitude

    watches the guardrail metrics during the run, so nobody has to peek at the primary result to stay informed

  • Optimizely

    runs the test itself against the sample size and stopping rule fixed before launch

Tell us what you want to test and which decision the result needs to support. We will set the rules before launch and apply them consistently, whatever the outcome.
Plan an experiment

Only if we planned for that upfront with a proper sequential testing method. Stopping an unplanned test the moment it looks good is one of the most common ways experiments produce false positives.