An experiment cannot be more trustworthy than its QA record.

A visual editor shows one browser and one happy path. It doesn't show the mobile keyboard covering the CTA, the tracking script that quietly stopped firing, or the accessibility regression a screen-reader user hits first. We run every variant through the same checklist, verify assignment and tracking before exposure, and exercise the rollback path in advance. A standing release gate makes sense once experiments run often enough that checks by feel create avoidable risk.

You get a signed QA record for every experiment before launch, plus a rollback path that has already passed a live drill.

A signed release checklist covering functional, tracking, performance, and accessibility QA for an experiment variant

Some of the 500+ brands we've worked with

See all references
  • GE
  • Yves Rocher
  • Peak Games
  • eOfis
  • Hotiç
  • Turna.com

We apply the approved checklist to each variant and keep the evidence with the release record. AI can draft checks and flag mismatches. The experiment owner is the person who approves release.

How we hold ourselves to it

  • Every variant clears one release checklist
  • Assignment logs must match the configured split before exposure begins
  • QA reviews performance and accessibility against the base page
  • The rollback mechanism is exercised in advance
  1. Build the release checklist for this program

    Before the first experiment enters QA, define the checks required for this site or app, including functional behavior, cross-device rendering, accessibility, performance, tracking, and assignment.

    Program release checklist covering functional, cross-device, accessibility, performance, tracking and assignment checks.

    AI assist
    Drafts a checklist template from the site's existing QA standards and past incident history.
    Human gate
    The CRO strategist and engineering owner approve the checklist before it's used.
    Owners
    CRO strategist + engineering owner
  2. Verify functional and cross-device behavior

    Check the paused variant across the browsers, viewports, and devices that matter for this audience. Confirm that the intended change renders and behaves as designed in the full matrix.

    Signed device and browser matrix result with every discrepancy noted.

    AI assist
    Runs the variant through the configured device and browser matrix and flags visual or functional discrepancies.
    Human gate
    An independent reviewer who did not build the variant signs off on functional QA.
    Owners
    Independent QA reviewer + engineering
  3. Verify tracking and assignment integrity

    Confirm every relevant event fires correctly in each variant and that the assignment mechanism splits traffic the way it was configured, before any traffic is exposed.

    Event-firing and assignment-split verification recorded per variant before exposure.

    AI assist
    Compares fired events and assignment logs against the experiment's configuration and flags mismatches.
    Human gate
    The analytics owner confirms tracking is correct before the experiment owner authorizes exposure.
    Owners
    Analytics owner + experiment owner
  4. Check performance and accessibility regressions

    Measure the variant's effect on load performance and layout stability. Re-run an accessibility check against the same criteria used for the base page before release.

    Performance and accessibility comparison of the variant against the control page.

    AI assist
    Runs performance and accessibility checks on the variant and flags any regression against the control.
    Human gate
    The engineering owner reviews flagged regressions before sign-off.
    Owners
    Engineering owner + accessibility reviewer
  5. Prove the rollback path works

    Exercise the rollback mechanism before the experiment goes live. The drill must leave a working exit if a guardrail breach occurs.

    Rollback drill record showing the exit worked before launch.

    AI assist
    Documents the rollback steps and flags any step that depends on a manual action with no assigned owner.
    Human gate
    The experiment owner confirms the rollback test passed before authorizing launch.
    Owners
    Experiment owner + engineering
  6. Sign, launch, and log the record

    Collect sign-off from every required reviewer, launch only the approved scope, and file the complete QA record so it's available if a guardrail trips or a post-mortem is needed later.

    Complete, signed QA record filed with the launch scope it approved.

    AI assist
    Compiles the signed checklist, test results, and rollback confirmation into one filed record.
    Human gate
    The experiment owner authorizes launch only once every required sign-off is on file.
    Owners
    Experiment owner + CRO strategist

These four records show which checks ran, what they found and who approved the variant before launch.

  • Program-level release checklist

    The standing checklist every experiment on this program gets checked against, covering functional, cross-device, tracking, performance, and accessibility QA.

    Accepted when

    Every required check for this site or app is written down once, and no experiment reaches launch without being measured against it.

    Cadence: Built once

  • Functional and cross-device QA record

    The signed result of running each variant through the configured device and browser matrix, with any discrepancies noted.

    Accepted when

    Each variant has been run through the configured device and browser matrix and every discrepancy is written down, not waived silently.

    Cadence: Before every launch

  • Tracking, assignment, performance, and accessibility findings

    The QA record confirms that events fired, assignment split as configured, and performance and accessibility were reviewed against the control before release.

    Accepted when

    Events fired in each variant, assignment split as configured, and performance and accessibility were compared against the control before release.

    Cadence: Before every launch

  • Tested rollback record

    The signed result of exercising the rollback mechanism and confirming it worked before launch.

    Accepted when

    The rollback mechanism was exercised, not just described, and the drill result is signed before exposure starts.

    Cadence: Before every launch

A variant that looks correct in the builder can still fail in production. The builder preview is only the starting point. Once variant code meets real browsers and devices, it can shift the layout, interrupt tracking, slow the page, or create an accessibility regression. QA has to inspect that production behavior before visitors are exposed.

A good fit when

  • Experiments launch often enough that checks by feel create avoidable release risk, yet no repeatable checklist shows what every variant must clear.
  • Variants come from an internal team or agency, but nobody independent has checked their tracking, device behavior, performance, and accessibility before exposure.
  • A failed experiment release would disrupt real traffic, so the team needs the rollback path tested before a guardrail ever trips.
  • One desktop browser is the entire review.
  • Assignment splits go unchecked until the results look odd.
  • A new script or style changes page load or layout stability before anyone measures it.
  • The team has no tested rollback path once exposure starts.

Better handled as other work when

  • You run only a handful of experiments each year. The preflight inside test design already covers them.
  • The variants still exist only as ideas, so there is no staged build, assignment log, or event firing to inspect yet.
  • QA keeps finding issues, but no one owns the release or rollback decision, so every flag waits without a route to action.

Paid Search, Paid Social, CRO, and Programmatic each run under a named owner at Zeo. The consultants below are matched to the channel this page is about, so you can see who you'd actually work with.

  • BrowserStack

    catches the mobile keyboard covering the CTA on a real device, not an emulator's guess at one

  • Google Tag Manager

    shows which tags actually fired, and in what order, before the variant goes live

We can build or run the release gate, catch tracking, performance, accessibility and functional problems before exposure, then exercise the rollback path.
Plan experiment QA

Test design includes a preflight step for that one experiment. This method is the standing, program-level discipline for teams running enough experiments that a repeatable, independently reviewed checklist matters more than a one-off check.