Retrieval, groundedness, citations, access rules, and abstention fail in different ways, so each carries a separate score before we judge the complete answer. A RAG release decision is credible on that basis.

A fluent RAG answer can rest on the wrong document, an unsupported claim, a citation that points nowhere useful, or a retrieval the reader's role never allowed. We score each of those separately, so a failure lands on the layer that owns it instead of blurring into one quality number.

At the release decision, every open failure sits with the RAG layer that owns it, backed by golden queries, subsystem grading, query-slice findings, and a release log.

Illustration of RAG Evaluation & Groundedness Testing: a team wiring a document pipeline into a retrieval system

Some of the 500+ brands we've worked with

See all references
  • Amazon
  • English Home
  • Yandex
  • Odeabank
  • İstanbul Gedik Üniversitesi
  • Desa
  • GS Store

One blended answer score can't tell you what to fix. We score each part of the RAG path on its own first, and only then judge how the complete answer behaves.

  1. Set the test boundary

    We lock in the query population, source corpus, access roles, expected evidence, known failures, and release thresholds.

    AI assist
    The model drafts a candidate query population from known usage and failures so the product owner can trim it.
    Human gate
    Does the evaluation set represent the questions and users that matter? Your product owner confirms the evaluation set represents the questions that matter.
  2. Build protected test cases

    We create golden queries and relevance judgments, then keep the evaluation set separate from implementation tuning.

    AI assist
    Candidate relevance judgments come pre-drafted, and a domain reviewer accepts or rejects each one.
    Human gate
    Can each judgment be traced to an approved source and review rule? A domain reviewer confirms each judgment before it enters the protected set.
  3. Test each subsystem

    We score retrieval, grounded claims, citations, access, and abstention separately before reviewing the complete response path.

    AI assist
    Per-case tooling scores retrieval, citations, and groundedness as separate lines, so no component hides inside a total.
    Human gate
    Is each failure assigned to the subsystem that can actually fix it? Your evaluation lead confirms which subsystem owns a given failure.
  4. Retest the release gate

    We rerun failed slices after changes and record open exceptions, regression status, and the next owner decision.

    AI assist
    Failed slices rerun automatically and arrive as a draft regression status a reviewer confirms.
    Human gate
    Does the product owner release, release with conditions, or hold the build? Your product owner releases, conditions, or holds the build.

Every finding stays attached to the query, source, role, and subsystem where it appeared, so the fix conversation starts in the right room.

  • Dataset

    Golden-query and relevance-judgment inventory

    Captures representative queries, expected sources, relevance labels, access context, and review rules.

  • Test evidence

    Retrieval-to-abstention grading matrix

    Tests retrieval, groundedness, citation, access, and abstention both separately and across the full response path.

  • Report

    Query-slice failures and remediation findings

    Groups errors by query type, source, role, subsystem, and attempted remediation.

  • Decision record

    RAG thresholds, exceptions, and release log

    Records thresholds, critical exceptions, approved conditions, owners, and the next regression trigger.

Retrieval, grounding, citation, access, and abstention fail in different ways and get fixed by different people. If your RAG metrics can't tell those failures apart, every remediation is a guess. That's the case this work is built for.

A good fit when

  • Your RAG answers look fluent but fail review, and the team cannot tell whether retrieval, unsupported generation, or citation caused the error.
  • Answer quality looks acceptable overall, while access leaks and failures to abstain remain hidden inside the aggregate result.
  • Your representative queries and source documents exist, but known failures are not tied to access roles, traces, and expected evidence.
  • Your golden queries are drafted, yet relevance judgments, critical user slices, and protected cases do not share one review rule.
  • The full response receives one score, so retrieval, groundedness, citation, access, and abstention failures remain indistinguishable.
  • A remediation lifts the average, but failed query slices are still missing from the regression gate.
  • Component evidence exists, while critical exceptions and regression status still leave the product owner unable to make a release call.

Better handled as other work when

  • You expect groundedness and abstention results to support certification or an audit opinion. Your product authority owns the legal and regulatory determination.
  • You need every answer to be correct, grounded, complete, or risk-free, although the evidence covers only the tested corpus, roles, and queries.
  • You need corpus acquisition, production operation, or remediation delivered, because this work stops at the agreed RAG test boundary.

If one of these is closer to your situation, start here instead: See the evaluation service

It's hard to test a system well if you've never had to keep one running. We operate production AI ourselves, so our evaluation, security testing, and LLMOps work starts from what actually breaks. The people on it are senior engineers, and Zeo has been doing client work since 2011.

  • LlamaIndex

    exposes the retrieval pipeline's intermediate steps as their own testable subsystem

  • Chroma

    the indexed corpus retrieval tests run against, matching what production queries

  • Ragas

    scores retrieval and groundedness separately, matching this page's subsystem split

  • Confident AI / DeepEval

    checks whether retrieval respects the reader's role, not just document relevance

Bring the current pipeline, the questions your users actually ask, and the person who has authority over release.
Talk to Zeo

The in-scope RAG implementation, representative queries and source documents, access roles, expected evidence, known failures, traces, and the release thresholds you're proposing. We also settle up front how sensitive documents and evaluation cases may be used and retained.