Diagrams assert that a boundary is isolated, logged, or recoverable. We follow each claim across identities, data, models, and tools for one agreed design state, then compare it with the evidence that exists.

Architecture diagrams often state that a boundary is isolated, logged, or recoverable. We follow those claims across identities, data, models, and tools, then compare them with the evidence available for one agreed design state. Open exceptions return with owners and the checks required for closure.

Open exceptions come back with owners and the check each one needs to close, and your architecture and risk owners decide which stay accepted and which are deferred.

Illustration of Secure AI Architecture Review: a team probing an AI system for security weaknesses

Some of the 500+ brands we've worked with

See all references
  • Shell
  • LC Waikiki
  • D&R
  • MNG Kargo
  • Sporx
  • Duru
  • Teyit.org

We fix the architecture version first, follow each important security claim to its available support, and leave every finding with an owner and a specific condition for another review.

  1. Fix one architecture state for review

    We assemble the current diagrams and records for trust boundaries, identities, permissions, data flows, model and tool access, isolation, logging, and recovery. Missing or conflicting material remains a review item instead of being reconciled by assumption.

    AI assist
    A model compares the approved diagrams with configuration records and flags where the two appear to diverge.
    Human gate
    Does the reviewed version contain the systems, dependencies, and owners your architecture owner expects? Your architecture owner confirms that the fixed view contains the systems consequential to the decision.
  2. Ask what supports each design claim

    Every material claim is read beside representative examples, current constraints, baseline evidence, dependencies, and known exceptions. The register identifies the source and states plainly where a conclusion still relies on an assumption.

    AI assist
    We use it to place approved evidence beside each architecture claim, leaving missing support visible.
    Human gate
    Does the available evidence support each consequential claim, or must it remain an assumption? Our review lead decides whether the evidence is sufficient for the claim or whether it stays unresolved.
  3. Work through credible design failures

    We examine representative cases involving access, isolation, logging, recovery, models, tools, and data flows. Each finding names the affected boundary, evidence considered, exception state, and decision it calls for.

    AI assist
    The approved architecture view provides the boundaries for draft failure cases that specialists refine.
    Human gate
    Which critical findings need remediation evidence, and which require a documented risk decision? A specialist confirms which findings warrant remediation or an explicit risk decision.
  4. Define what closure will look like

    Every open item receives an owner, acceptance condition, and the evidence or test needed for another decision. Your architecture and risk owners then accept, reject, or defer the finding.

    AI assist
    Agreed closure conditions are turned into a first retest plan for the owners to review.
    Human gate
    For every open finding, are the owner, decision state, and specific closure evidence recorded? Your architecture and risk owners make the accept, reject, or defer decision.

Reviewers can move from a finding to the evidence behind it, see which assumption or exception remains open, and identify exactly what has to happen before closure.

  • Architecture document

    Security architecture findings and retest plan

    Findings across trust boundaries, identities, data flows, model and tool access, isolation, logging, and recovery, with the retest route for open items.

  • Risk register

    Reviewed sources and the claims still waiting on evidence

    Reviewed sources, explicit assumptions, system dependencies, current constraints, and questions still waiting on evidence.

  • Report

    Failure-case notes and the gaps they exposed

    Representative failure cases, the evidence considered, exceptions raised, and critical gaps taken into the acceptance discussion.

  • Decision record

    Closure brief for the accepted and deferred findings

    Accepted and deferred findings, the people responsible, remaining exposure, and the next review or stop point.

Bring this review in before an AI design reaches production or after a live design changes how models, data, identities, or tools interact.

A good fit when

  • Two diagrams of the same system disagree, and the people who could settle which one is current sit in different teams.
  • Model or tool access changed after the last review, so isolation, logging, and recovery have not been examined as one system since.
  • Someone can halt the release, but they are being asked to decide on a diagram rather than on evidence that the boundary holds.
  • Trust boundaries, identities, permissions, data flows, and model or tool access were each drawn by a different team, and none of them meet on one page.
  • A design claims isolation, logging, and recovery, yet nobody has asked what evidence supports each claim for the version running now.
  • Credible failure cases have never been walked through the design, so acceptance conditions and exception owners are still implicit.
  • A finding will need a defined closure condition, though today nothing says what evidence or control change would let it be reviewed again.

Better handled as other work when

  • You need the architecture signed off as compliant. We report what the evidence supports for one design state, and that call belongs to your qualified authority.
  • You want a guarantee that the design is safe or that no residual risk remains. A review reads evidence against claims and cannot promise either.
  • You want the findings fixed and the system run for you. Remediation and production operation are separate commissions from this review.

If one of these is closer to your situation, start here instead: See security review options

It's hard to test a system well if you've never had to keep one running. We operate production AI ourselves, so our evaluation, security testing, and LLMOps work starts from what actually breaks. The people on it are senior engineers, and Zeo has been doing client work since 2011.

  • Amazon Web Services

    the account configuration a boundary claim on AWS-hosted AI gets checked against

  • Microsoft Azure AI

    the Azure-side configuration a boundary claim on Azure-hosted AI gets checked against

  • Datadog

    the alert and log history a 'this is logged' claim gets checked against

  • IriusRisk

    generates a diagram from text, then derives the threat model from it

  • Lakera Guard

    the runtime enforcement point a claimed guardrail layer gets probed against

  • Guardrails AI

    the schema check run at the exact point a diagram marks as an output boundary

Bring the current system view, known exceptions, and the owners authorized to decide what changes or remains open. Each consequential claim is then tested against the evidence behind it.
Inspect the architecture

Start with the current trust boundaries, identities, permissions, data flows, model and tool access, isolation, logging, and recovery. We also need the owners, representative examples, constraints, dependencies, and baseline evidence behind that view. Incomplete or conflicting material can stay in the review if its status is clear.