An agent can look bounded in its prompt and still hold a tool permission that changes the outcome. Identities, MCP connections, tool contracts, grants, instruction boundaries, and side effects have to be tested as one system.

An agent may look bounded in a prompt and still hold a tool permission that changes the outcome. We test identities, MCP connections, tool contracts, grants, instruction boundaries, and side effects as one system. Reproducible findings return to your security owner for exception and release decisions.

Your security owner decides exceptions and release from findings that reproduce on demand, each traced back to the grant that permitted it and retested after the fix.

Illustration of Agent, Tool & MCP Security Testing: a team probing an AI system for security weaknesses

Some of the 500+ brands we've worked with

See all references
  • Abdi İbrahim
  • Millenicom
  • Shiftdelete
  • Sina Pırlanta
  • Armut.com
  • S Sport Plus

We first make the agent's authority visible. Approved paths are then exercised with retained evidence, and material exceptions return to the same boundary after remediation.

  1. Map the action path

    We inventory agent and MCP identities, tool contracts, permissions, instruction boundaries, expected side effects, known exploit paths, dependencies, and accountable owners. The map also records components and actions that are outside the authorization.

    AI assist
    Approved configurations and permission grants are compared in a draft boundary inventory. A specialist checks every entry.
    Human gate
    Can your security owner rely on the mapped identities, tools, permissions, and effects for the release decision? Your security owner authorizes the identities, tools, effects, and paths included in testing.
  2. Reproduce the approved exploit paths

    We run representative cases through the mapped boundary and retain the inputs, permissions, outputs, and side effects needed to explain what the agent, MCP component, or tool did. Stop conditions remain active throughout.

    AI assist
    Known patterns and approved tool information seed bounded variations of the authorized cases.
    Human gate
    Can an authorized second tester reproduce the observed behavior from the same inputs and grants? Your security lead approves each exploit path before it is executed and can stop the test.
  3. Link the exception to its cause

    Results are compared with the agreed acceptance conditions. Each critical exception stays attached to the relevant identity, contract, permission, prompt boundary, tool call, or side effect so the owner can see what must change.

    AI assist
    Observed results are grouped against the acceptance conditions, with possible critical exceptions flagged for review.
    Human gate
    Which verified exceptions block acceptance, and which require an explicit residual-risk decision? The specialist you appoint decides which exceptions block acceptance.
  4. Retest only what changed

    After remediation, we rerun the affected authorized paths and place the new evidence beside the original finding. The handoff records the accepted result, residual risk, accountable owner, and next review point.

    AI assist
    Approved retest evidence is formatted into a draft summary and decision record for the owner.
    Human gate
    Has the owner accepted the remaining risk and taken responsibility for the next review or stop point? Your accountable owner accepts the residual risk and signs the handoff.

The evidence package lets an authorized reviewer see what the agent was allowed to do, reproduce a finding, and identify the person responsible for every open exception.

  • Test evidence

    Agent-tool security findings and their retests

    Findings and retests organized around the authorized identity, MCP component, tool contract, permission, instruction boundary, side effect, and exploit route.

  • Risk register

    Baseline evidence and the questions left open

    Approved inputs, assumptions, dependencies, current constraints, baseline evidence, and questions that remain unresolved.

  • Report

    Exercised-path log and the critical exceptions

    The cases exercised, their acceptance results, critical exceptions, and the identities, tools, permissions, or effects involved.

  • Decision record

    Acceptance note with the open risks and their owners

    The accepted result and conditions, plus every open risk, its owner, and the next review or stop point.

Bring us in when an agent can reach data, tools, or external systems through permissions and MCP connections without one reviewed action boundary.

A good fit when

  • Platform, security, and the agent team each grant part of the permission set, so no single person can say what the agent is allowed to reach.
  • Someone has demonstrated a worrying agent behavior once, but nobody has reproduced it on demand or traced it back to the grant that allowed it.
  • Release is blocked on a security answer, and the person who can approve an exception wants evidence, not an opinion.
  • Identities, MCP connections, tool contracts, prompt boundaries, and side effects were each reviewed separately, so nobody has tested them as one path.
  • Constraints, dependencies, and known exceptions live in different heads, while nobody has collected baseline evidence for the agent's current grants.
  • An exception was accepted once, but there is no retest that would tell you whether the fix held or the exception simply stopped being mentioned.
  • Coverage is asserted, though nobody can point to which identity, MCP component, or tool contract was left outside the tested boundary.

Better handled as other work when

  • You need a security certificate or an audit opinion out of this. The report carries reproducible findings, and those judgments stay with the authority you appoint.
  • You want a statement that every future agent action and tool behavior is safe. Testing covers the grants and paths that existed on the day we ran it.
  • You want the fixes implemented and the agent run for you. Both sit outside the authorized test boundary unless commissioned separately.

If one of these is closer to your situation, start here instead: Explore security consulting

It's hard to test a system well if you've never had to keep one running. We operate production AI ourselves, so our evaluation, security testing, and LLMOps work starts from what actually breaks. The people on it are senior engineers, and Zeo has been doing client work since 2011.

  • MCP-Scan

    the MCP-specific scanner reading tool descriptions for hidden injection payloads

  • Mindgard

    the recon-to-runtime campaign covering an agent's own tool-calling surface continuously

  • Lakera Guard

    the runtime layer enforcing a tested agent boundary continuously in production

  • Giskard

    the continuous scan covering prompt injection, data disclosure, and sycophancy risks

Share the agent, MCP, and tool map with the owner who can approve the test and act on the result. The review stays inside that authority.
Scope the test

Which inputs have to be ready before testing?

We need the agent and MCP identities, tool contracts, permission grants, prompt boundaries, representative examples, constraints, baseline evidence, and whoever is authorized to approve the test and accept its result.

Will you test every possible agent action?

No. We agree a representative boundary and make its coverage visible. The result applies to the paths, permissions, and effects exercised in that system state. A later tool call, grant, or exploit route needs its own evidence.

What supports the release decision?

We look at the agreed acceptance results and the closure state of each critical exception. Evidence must cover the relevant identities, contracts, permissions, prompt boundaries, effects, and exploit paths. The handoff also needs a named owner and decision state. Each critical exception is judged on its own, regardless of the overall pass count.

What happens after a path fails?

We preserve enough evidence to reproduce the behavior, identify the affected boundary, and route the exception to its owner. After remediation, we repeat the authorized path. Evidence informs the decision, but a person still decides how much residual risk the organization will carry.