Adversarial campaign · AI red teaming
AI Red Teaming
An AI red-team campaign stays useful only when every attack path is authorized, every finding reproduces, and the same bounded path faces retesting after remediation.
We test how an adversary could combine weaknesses across the model, application, data, tools, users, and operations. You approve the campaign and its stop conditions first. Findings must reproduce before they enter the risk record, and agreed fixes face the same paths again.
Verified attack-path findings, severity judgments, and a fixed-path retest pack let your authority separate closed exposure from accepted residual risk.


Some of the 500+ brands we've worked with
See all referencesSteps, gates, and who decides
How we work
Every stage leaves evidence a person can inspect. Authorization, severity, exception handling, release, and residual-risk acceptance remain human decisions from start to finish.
Sign the rules of engagement
Before a live system is touched, we agree the campaign objectives, in-scope paths, exclusions, evidence handling, safety controls, access, and stop conditions. The attack matrix records what is permitted and who can halt the work.
- AI assist
- Candidate attack-matrix entries are extracted from the approved threat model. The security lead approves or removes them.
- Human gate
- Have the attack plan, authorized targets, handling rules, and stop authority all been approved in writing? Your security lead approves the attack paths and can stop the campaign.


Exercise the campaign across system layers
We run manual and automated tests against the model, application, data, and tools inside the approved matrix. A path ends when the rules of engagement say it ends, even if further exploration appears possible.
- AI assist
- Approved sources can be organized and case variants drafted inside the authorized workspace. A specialist reviews both.
- Human gate
- Does the observed behavior reproduce under the authorized conditions and retained evidence? A Zeo specialist verifies each reproduced result before it counts as a finding.


Validate the finding before rating it
A candidate reaches the risk report only after reproduction, evidence preservation, and review of severity and exploit-chain depth against the thresholds agreed before testing.
- AI assist
- Recorded exploit-chain evidence supports a draft severity view. A person makes the final rating.
- Human gate
- Where do the agreed thresholds place each verified finding: pass, pass with conditions, or fail? The person you appoint makes the final severity and acceptance decision for each finding.


Retest the fixed paths
After remediation, we rerun the same authorized paths and issue the decision record. Your named owner states what closed, what remains open, and which exposure is being accepted as residual risk.
- AI assist
- Approved retest evidence and workshop notes become a first decision summary for the owner.
- Human gate
- Has the owner accepted the bounded retest result, residual risk, and date or event for another review? Your owner accepts the result and decides what remains residual risk.


Named artifacts you keep
What you get
The campaign leaves a versioned account of what was authorized, what ran, what reproduced, how it was judged, and who made the final decisions.


Matrix
Authorized campaign rules and attack-path charter
Approved objectives, systems and paths in scope, exclusions, safety controls, stop conditions, evidence rules, and planned tests.


Test evidence
Reproduction steps and bounded technical findings
Each verified finding includes bounded evidence and reproduction steps that an authorized engineer can use to check the result.


Risk register
Validated findings reports for committee and engineers
The same validated findings are presented at the level needed by the risk committee and the engineers responsible for remediation.


Workshop record
Remediation choices and fixed-path retest pack
Workshop decisions and retest results show which agreed paths closed, which remain open, and which were outside the retest.
Scope and honest limits
When to bring us in
Red teaming starts with written authorization. Targets, safety limits, stop conditions, evidence handling, and final authority must be settled before live testing.
A good fit when
- Verified findings reach the risk record, but nobody with authority is named to decide which residual exposure the organization will carry.
- Your threat model and test environment exist, but no authorized adversary has exercised the in-scope system and preserved campaign evidence.
- A campaign starts, but severity thresholds, exclusions, stop conditions, and exception rules remain unwritten.
- Live systems need authorized testing, but the rules of engagement, campaign paths, and safety controls are still scattered or unwritten.
- The model, application, data, and tool weaknesses have been reviewed separately, while no approved attack matrix connects them as one path.
- Candidate findings exist, but nobody has reproduced them, rated severity, or agreed which authorized paths must be retested after remediation.
- Scope changes and stop events happen during testing, yet the human decision gates and residual-risk trail cannot be reconstructed afterward.
Better handled as other work when
- You need the campaign to issue legal, regulatory, or certification sign-off. Your counsel and compliance leads issue those judgments from the campaign evidence.
- You need a blanket claim that the system is safe, although the attack matrix covers only the paths and conditions you authorized.
- You expect every finding to be fixed during the campaign, but remediation happens only when scoped beyond verification and retesting.
If one of these is closer to your situation, start here instead: View AI security services
We operate the systems we test
It's hard to test a system well if you've never had to keep one running. We operate production AI ourselves, so our evaluation, security testing, and LLMOps work starts from what actually breaks. The people on it are senior engineers, and Zeo has been doing client work since 2011.
Tools we use
Tools behind this work
Promptfoothe open-source probe generator producing a quantified risk report before and after a fix
Mindgardthe continuous cross-layer campaign engine an adversary's combined path gets exercised through
garakthe NVIDIA-backed LLM probe library supplying broad coverage before custom work starts
Lakera Guardthe self-serve automated workflow a campaign's approved scope runs against continuously
Next step
Set the campaign rules before the first attack


Before you decide




























