Agent security review · Tools and MCP
Agent, Tool & MCP Security Testing
An agent can look bounded in its prompt and still hold a tool permission that changes the outcome. Identities, MCP connections, tool contracts, grants, instruction boundaries, and side effects have to be tested as one system.
An agent may look bounded in a prompt and still hold a tool permission that changes the outcome. We test identities, MCP connections, tool contracts, grants, instruction boundaries, and side effects as one system. Reproducible findings return to your security owner for exception and release decisions.
Your security owner decides exceptions and release from findings that reproduce on demand, each traced back to the grant that permitted it and retested after the fix.


Some of the 500+ brands we've worked with
See all referencesSteps, gates, and who decides
How we work
We first make the agent's authority visible. Approved paths are then exercised with retained evidence, and material exceptions return to the same boundary after remediation.
Map the action path
We inventory agent and MCP identities, tool contracts, permissions, instruction boundaries, expected side effects, known exploit paths, dependencies, and accountable owners. The map also records components and actions that are outside the authorization.
- AI assist
- Approved configurations and permission grants are compared in a draft boundary inventory. A specialist checks every entry.
- Human gate
- Can your security owner rely on the mapped identities, tools, permissions, and effects for the release decision? Your security owner authorizes the identities, tools, effects, and paths included in testing.


Reproduce the approved exploit paths
We run representative cases through the mapped boundary and retain the inputs, permissions, outputs, and side effects needed to explain what the agent, MCP component, or tool did. Stop conditions remain active throughout.
- AI assist
- Known patterns and approved tool information seed bounded variations of the authorized cases.
- Human gate
- Can an authorized second tester reproduce the observed behavior from the same inputs and grants? Your security lead approves each exploit path before it is executed and can stop the test.


Link the exception to its cause
Results are compared with the agreed acceptance conditions. Each critical exception stays attached to the relevant identity, contract, permission, prompt boundary, tool call, or side effect so the owner can see what must change.
- AI assist
- Observed results are grouped against the acceptance conditions, with possible critical exceptions flagged for review.
- Human gate
- Which verified exceptions block acceptance, and which require an explicit residual-risk decision? The specialist you appoint decides which exceptions block acceptance.


Retest only what changed
After remediation, we rerun the affected authorized paths and place the new evidence beside the original finding. The handoff records the accepted result, residual risk, accountable owner, and next review point.
- AI assist
- Approved retest evidence is formatted into a draft summary and decision record for the owner.
- Human gate
- Has the owner accepted the remaining risk and taken responsibility for the next review or stop point? Your accountable owner accepts the residual risk and signs the handoff.


Named artifacts you keep
What you get
The evidence package lets an authorized reviewer see what the agent was allowed to do, reproduce a finding, and identify the person responsible for every open exception.


Test evidence
Agent-tool security findings and their retests
Findings and retests organized around the authorized identity, MCP component, tool contract, permission, instruction boundary, side effect, and exploit route.


Risk register
Baseline evidence and the questions left open
Approved inputs, assumptions, dependencies, current constraints, baseline evidence, and questions that remain unresolved.


Report
Exercised-path log and the critical exceptions
The cases exercised, their acceptance results, critical exceptions, and the identities, tools, permissions, or effects involved.


Decision record
Acceptance note with the open risks and their owners
The accepted result and conditions, plus every open risk, its owner, and the next review or stop point.
Scope and honest limits
When to bring us in
Bring us in when an agent can reach data, tools, or external systems through permissions and MCP connections without one reviewed action boundary.
A good fit when
- Platform, security, and the agent team each grant part of the permission set, so no single person can say what the agent is allowed to reach.
- Someone has demonstrated a worrying agent behavior once, but nobody has reproduced it on demand or traced it back to the grant that allowed it.
- Release is blocked on a security answer, and the person who can approve an exception wants evidence, not an opinion.
- Identities, MCP connections, tool contracts, prompt boundaries, and side effects were each reviewed separately, so nobody has tested them as one path.
- Constraints, dependencies, and known exceptions live in different heads, while nobody has collected baseline evidence for the agent's current grants.
- An exception was accepted once, but there is no retest that would tell you whether the fix held or the exception simply stopped being mentioned.
- Coverage is asserted, though nobody can point to which identity, MCP component, or tool contract was left outside the tested boundary.
Better handled as other work when
- You need a security certificate or an audit opinion out of this. The report carries reproducible findings, and those judgments stay with the authority you appoint.
- You want a statement that every future agent action and tool behavior is safe. Testing covers the grants and paths that existed on the day we ran it.
- You want the fixes implemented and the agent run for you. Both sit outside the authorized test boundary unless commissioned separately.
If one of these is closer to your situation, start here instead: Explore security consulting
We operate the systems we test
It's hard to test a system well if you've never had to keep one running. We operate production AI ourselves, so our evaluation, security testing, and LLMOps work starts from what actually breaks. The people on it are senior engineers, and Zeo has been doing client work since 2011.
Tools we use
Tools behind this work
MCP-Scanthe MCP-specific scanner reading tool descriptions for hidden injection payloads
Mindgardthe recon-to-runtime campaign covering an agent's own tool-calling surface continuously
Lakera Guardthe runtime layer enforcing a tested agent boundary continuously in production
Giskardthe continuous scan covering prompt injection, data disclosure, and sycophancy risks
Next step
Test the authority behind the agent


Before you decide


























