AI Evaluation & Assurance · Independent agent assurance
Independent AI Agent Evaluation
An agent review is independent only when the build team controls neither test scope nor reporting, and material findings reproduce under the evaluator's own tools and trace review.
The team that built an agent grades its own homework more kindly than anyone else will. Under a documented independence charter, we test the existing agent across tasks, trajectories, tools, policies, and recovery, with scope and conclusions the implementation team doesn't control.
The independent recommendation covers the tested paths and nothing beyond them. Your release owner can read the chartered outside test record, the reproduced findings, the separated implementation claims, and the residual risk.


Some of the 500+ brands we've worked with
See all referencesSteps, gates, and who decides
How we work
Independence fails quietly when scope drifts mid-review. We pin down scope, access, reporting lines, and evidence provenance first, so the recommendation can stand apart from implementation self-testing.
Set the independence charter
We define the test scope, access boundaries, reporting lines, and what counts as independent versus implementation-provided evidence.
- AI assist
- The client's brief becomes a draft scope and access boundary the engagement owner corrects.
- Human gate
- Is the charter specific enough to keep the review genuinely independent? Your engagement owner confirms the charter keeps the review genuinely independent.


Run and reproduce
We repeat representative and adverse tasks, deterministic checks, and calibrated rubric review across multiple runs.
- AI assist
- Repeated trials run under our own tooling, which logs the raw results for review.
- Human gate
- Does each material result hold up on repeated, independently run trials? Our evaluation lead confirms a result before it's treated as reproducible.


Attribute variance and exploits
We separate normal run-to-run variance from genuine exploits, slice failures, and recovery gaps, and check severity against agreed criteria.
- AI assist
- Run-to-run variance gets clustered apart from candidate exploits so reviewers don't conflate the two.
- Human gate
- Is each finding backed by a reviewed trace rather than a single anomalous run? A risk specialist decides whether a flagged item is a genuine exploit.


Issue the independent recommendation
We retest disputed or fixed paths and hand over a recommendation kept separate from the implementation team's own claims.
- AI assist
- The independently run evidence becomes a draft recommendation summary the independent lead edits and signs.
- Human gate
- What does the independent evidence support: release, conditional release, or hold? Our independent lead issues the recommendation without input from the implementation team.


Named artifacts you keep
What you get
Every artifact preserves which evidence we produced ourselves and which claims came from the implementation team. That separation is the product.


Playbook
Independent scope, access, and reporting-line charter
Documents the review boundary, access terms, and what independence means for this engagement.


Decision record
Versioned agent task-and-trajectory provenance file
Captures test design, execution history, and the chain from result back to the exact agent version reviewed.


Test evidence
Trace-reviewed exploit and variance findings
Shows where deterministic checks, rubric graders, and human review agreed or diverged on each finding.


Decision record
Independent release recommendation and retest brief
States the independent recommendation, residual risk, and retest evidence for previously flagged paths.
Scope and honest limits
When to bring us in
Choose this when a high-stakes release, audit, customer request, or dispute needs evidence produced by a team that did not build the agent.
A good fit when
- The build team has its own test results, but a high-stakes release still lacks evidence produced under an independent scope and reporting line.
- A stakeholder, auditor, or customer needs outside evidence, yet every current claim comes from the implementation team's own trials.
- Agent access and prior traces exist, but the implementation team still controls which tasks are tested and which conclusions reach the release owner.
- Tasks, failures, trajectories, policies, and tools are under review, but no independence charter fixes the scope before testing starts.
- The material results vary across runs, while no independently operated trial and trace review has shown which findings reproduce.
- Ordinary variance, exploits, slice failures, and recovery gaps appear together, so a single anomalous run can be mistaken for a finding.
- A release recommendation is expected, but implementation evidence and independently produced results still share one conclusion and reporting path.
Better handled as other work when
- You need legal, regulatory, audit, or certification approval. The independent report states tested behavior, while formal approval stays with your authority.
- You want the same team to evaluate, fix, and operate the agent. Those duties need separate scopes because combining them breaks the review boundary.
- You want the outside verdict to replace your release decision. The recommendation supplies reproducible evidence, while accountability stays with your release owner.
If one of these is closer to your situation, start here instead: See the evaluation service
We operate the systems we test
It's hard to test a system well if you've never had to keep one running. We operate production AI ourselves, so our evaluation, security testing, and LLMOps work starts from what actually breaks. The people on it are senior engineers, and Zeo has been doing client work since 2011.
Tools we use
Tools behind this work
Braintrustruns the independent test suite under scope the review team controls, not the builders
Langfusecaptures each run's trajectory in enough detail to reproduce results exactly
Confident AI / DeepEvalscores against rubrics the review owns, kept separate from the builders' own
Mindgardattributes a discovered exploit to its specific cause, not just its symptom
Next step
Get an answer the build team can't grade


Before you decide


























