Each sampled answer becomes a claim-level record of what is supported, contradicted, omitted, dated incorrectly or still unresolved.

Sampled answers stay attached to their full context while the checkable claims and their source passages are pulled out. Trained reviewers decide support, materiality, severity and the appropriate escalation path. Brand, communications and legal-adjacent leads deciding whether an AI answer contains a material error that warrants action.

For the documented sample, you get a claim-level scorecard that separates supported, contradicted, outdated, disputed and unverifiable brand facts.

Figure checking an AI answer card against a fact ledger and raising a flag on a mismatch

Some of the 500+ brands we've worked with

See all references
  • Hyundai
  • Bayer
  • Watsons
  • Vitra
  • QNB Finansfaktoring
  • Hotiç
  • Elele

The full answer remains available while its claims and source passages are reviewed. Materiality, dispute status, severity and escalation all pass through named human owners.

How we hold ourselves to it

  • Atomic claims over vibes
  • Tone judged apart from facts
  • Criticism stays if supported
  • Severity before escalation
  1. Define the claim and sentiment rubric

    The rubric names the canonical fact set, disputed fields, materiality rules, sentiment labels and the cases that must escalate.

    Fact and framing rubric with examples, exclusions and escalation thresholds.

    AI assist
    An agent organizes approved facts, disputed fields and worked materiality examples into the draft review rubric.
    Human gate
    Fact, legal and measurement owners approve the reference set and escalation boundaries.
  2. Decompose answers into atomic claims

    Complete answers stay intact while each fact, qualification, comparison and opinion receives its own reviewable record.

    Answer-claim corpus linked to full run context.

    AI assist
    The agent captures full responses, proposes the segments and keeps every segment linked to its original context.
    Human gate
    A reviewer validates segmentation samples and removes extraction artifacts.
  3. Match claims to trusted evidence

    Each claim is checked against an approved source, then assigned one of the rubric's evidence states.

    Claim-to-source matrix with evidence and reviewer rationale.

    AI assist
    Retrieves exact source passages and proposes support labels with dates and ownership.
    Human gate
    The accountable fact owner approves truth status or marks the claim disputed.
  4. Run calibrated automated classification

    The calibrated run flags omissions or framing that may materially change interpretation, while legitimate criticism and uncertainty remain in the record.

    Material omission and framing register with severity.

    AI assist
    The classifier applies the approved rubric and abstains where the evidence cannot support a label.
    Human gate
    Calibrated reviewers settle materiality and sentiment labels, including abstentions.
  5. Escalate material cases to human review

    Validated issues enter the queue according to harm, persistence, audience, source pattern and the owner able to act.

    Escalation queue with correction path and response deadline.

    AI assist
    Ranks validated issues by harm, persistence, audience, source pattern and reversibility.
    Human gate
    Legal, safety, crisis and business owners approve priority and response path in their remit.
  6. Trace likely sources and retest corrections

    Approved source changes trigger a replay of the affected prompts, plus cases withheld from the original run, under the original closure rule.

    Before/after accuracy report with remaining variance and closure evidence.

    AI assist
    Replays affected and held-out prompts and compares atomic claims against the original fact ledger.
    Human gate
    An independent reviewer closes, reopens or escalates the issue under the original rule.

Reviewers receive every material claim with its status, supporting evidence, accountable owner and closure test.

  • Answer accuracy scorecard

    A sampled claim ledger assigns every answerable brand fact its approved evidence state.

    Accepted when

    Correct owned sources, investigate propagation or accept uncertainty where truth is unresolved.

    Cadence: Claim-level evidence

  • Material misrepresentation report

    The material errors, omissions and misleading frames ranked by harm, persistence and audience.

    Accepted when

    Escalate now, plan a correction, monitor or document why no action is warranted.

    Cadence: Ranked by harm

  • Source correction priority map

    Exact source, page or fact-governance changes tied to each remediable issue and its approval owner.

    Accepted when

    Approve a bounded correction without manufacturing positive claims or suppressing supported criticism.

    Cadence: Owner assigned

  • Human-reviewed retest pack

    Post-change evidence from the repeated prompts and the withheld cases, showing which representation issues closed, which persisted, and which are new.

    Accepted when

    Close, reopen, broaden the diagnosis or roll back a harmful source change.

    Cadence: Human retested

  • Supported atomic-fact rate

    Use unsupported cases to correct sources or investigate. Do not average disputed facts into certainty.

    Accepted when

    Atomic answer facts fully supported by an approved current source divided by all answerable atomic facts in the audited sample.

    Cadence: Denominator disclosed

A sampled answer belongs here when a fact may be wrong or stale, a necessary condition is missing, or otherwise supported wording creates a materially misleading impression. A named person makes the decision in every case.

A good fit when

  • AI answers repeat wrong or stale facts — Repeated samples show materially incorrect product, company, people, location or policy details.
  • The wording is true but changes the likely reading — A missing condition or unbalanced comparison makes the supported claim materially misleading.
  • An automated sentiment score is driving action — No source comparison, reviewer agreement or escalation threshold makes the label usable.
  • One answer bundles several checkable facts — Names, dates, features, eligibility, prices, locations and comparisons cannot share one accuracy verdict.
  • The claim set mixes errors with unresolved facts — Approved sources cannot yet separate supported, outdated, disputed and unverifiable statements.
  • Tone is changing how a supported fact is read — A missing qualification or comparison matters enough to alter the reader's decision.
  • Tone and factual accuracy share one verdict — The rubric cannot separate positive, negative, mixed or uncertain framing from source support.

Better handled as other work when

  • You want supported criticism treated as an error — It keeps that status, and we do not pressure publishers or invent positive evidence for sentiment.
  • You need an automated label to settle truth — Sentiment and factual support need a calibrated human sample and an uncertain state.
  • Profound

    keeps the full context attached to a sampled answer, not just the flagged line

  • Google Search Console

    confirms the source page behind a claim is actually indexed and current

  • Airtable

    the reviewer's structured record, one row per claim with its own severity and escalation path

Sampled answers and the approved sources for disputed facts are enough to begin. We'll build the claim record, isolate material errors and show which cases are ready for correction or retest.
Check disputed AI claims

We keep the full answer, then split it into checkable claims. Each claim is compared with an approved, dated source. The result may be supported, contradicted, outdated, unverifiable or disputed. An automated confidence score does not settle truth.