We evaluate a safety policy in pairs. Each material rule faces a harmful behavior and a legitimate request that resembles it, because a system that blocks both has failed as surely as one that allows both.

A written safety rule is a claim until someone watches the system follow it. We turn each material rule into paired tests, a misuse attempt and a legitimate lookalike, then examine refusal quality and severity, retest what gets fixed, and record the risk that remains.

The decision package shows rule coverage, paired-scenario findings, violation severity, false refusals, retest results, and the residual risk your owner must address.

Illustration of Safety & Policy Evaluation: a team testing an AI system against representative evidence

Some of the 500+ brands we've worked with

See all references
  • MediaMarkt
  • PWC Türkiye
  • Sabancı Üniversitesi
  • Edenred
  • TRT
  • Cheetos
  • DYO

Every material rule gets the same treatment: a harmful case and a legitimate boundary case built around it, then severity review and retest for whatever fails.

  1. Trace each policy rule into behavior

    For every material rule, we record the intended behavior, the harm it addresses, plausible misuse, and a legitimate request that could be caught by the same control.

    AI assist
    The written policy text becomes a draft set of candidate test cases for the policy owner to correct.
    Human gate
    Can every material rule be traced to at least one observable test? Your policy owner confirms each rule's test matches its intent.
  2. Test both sides of the boundary

    We run each misuse attempt beside its legitimate lookalike. Deterministic checks capture clear outcomes, while calibrated reviewers examine refusals that depend on context or wording.

    AI assist
    The adversarial and benign scenario pairs run under tooling that logs every outcome for review.
    Human gate
    Can reviewers tell a real violation from a benign edge case using the same rubric? A reviewer confirms a flagged case is a real violation or a genuine benign edge case.
  3. Judge severity and refusal quality

    Reviewers classify what happened, how serious it is, which policy area it touches, and who is affected. Violations and false refusals remain separate throughout the analysis.

    AI assist
    Failures arrive pre-classified by severity and policy area so reviewers spend their time on judgment, not sorting.
    Human gate
    What do the agreed severity rules require for this policy area and affected slice? Your risk owner decides whether the severity thresholds pass, pass with conditions, or fail.
  4. Retest the fix and name what remains

    After a fix, we rerun the same paired cases and compare the evidence. The record states which behavior changed, which finding stayed open, and who is accountable for the remaining risk.

    AI assist
    Retest results come back as a draft residual-risk record the risk owner reviews.
    Human gate
    Does your risk owner accept, escalate, or send the residual risk back for more work? Your risk owner accepts, escalates, or sends the residual risk back for more work.

Every finding keeps its route from policy wording to observed behavior intact, so a disputed result can be re-checked line by line.

  • Matrix

    Rule-to-behavior test coverage map

    Maps each policy rule to its expected behavior, harm category, misuse pattern, and benign counterexample.

  • Test evidence

    Paired misuse and benign scenario review pack

    Contains calibrated adversarial and benign scenarios with expected outcomes and reviewer notes.

  • Report

    Paired-test violation, refusal, and severity findings

    Separates genuine policy violations from over-refusals and breaks results out by severity and affected group.

  • Decision record

    Remediation and paired-retest risk dossier

    Documents fixes attempted, retest results, and the residual risk your risk owner is accepting or escalating.

The failures that matter live near the policy boundary. A system may permit a harmful request and then reject a legitimate one written in similar language. We test those cases as pairs so neither failure hides.

A good fit when

  • Written policies exist but no one has tested whether the system actually follows them under pressure.
  • Benign requests are refused while similar harmful ones pass, so one overall rate hides both the business cost and the safety failure.
  • Your risk owner sees a single pass rate, but it does not show which severity, exception, or residual-risk decision needs attention.
  • Policy rules describe harm and expected behavior, but they have not been traced to paired misuse and benign edge-case tests.
  • Scenario outcomes differ with wording and context, so deterministic checks alone cannot settle them without a calibrated human rubric.
  • The review counts violations and false refusals together, which hides severity, affected slices, remediation status, and what remains after retest.
  • A fix has been retested, but your risk owner still lacks one residual-risk record that shows what changed, what remains, and who decides.

Better handled as other work when

  • The policy evaluation must serve as an audit opinion, legal or regulatory determination, or certification approval. Your qualified reviewers own those decisions.
  • You need a claim that the system is fully safe, compliant, or immune to future misuse. The paired tests only report behavior inside the reviewed scenario set.
  • You want the underlying policy rewritten before its rules can be tested. Policy design is separate work unless we agree to include it.

If one of these is closer to your situation, start here instead: See the evaluation service

It's hard to test a system well if you've never had to keep one running. We operate production AI ourselves, so our evaluation, security testing, and LLMOps work starts from what actually breaks. The people on it are senior engineers, and Zeo has been doing client work since 2011.

  • Promptfoo

    generates paired misuse and legitimate-lookalike tests straight from a policy rule

  • Mindgard

    runs adversarial attack variations a human tester wouldn't think to script

  • Lakera Guard

    classifies refusal quality and severity on real traffic after a fix ships

  • Giskard

    tracks residual risk across every policy rule as the structured end deliverable

Bring the written rules, the system behavior they govern, and the risk owner who can decide what happens after retesting.
Talk to Zeo

The intended use, applicable internal policies, system and tool design, misuse and harm concerns, representative users, current controls, and the person who owns risk decisions.