

Generative Engine Optimization
GEO Checklist
Verified 23 July 2026: audit 449 controls across 30 domains for AI-search eligibility, evidence quality, observability, and operational readiness. Every check includes What, How, Why, and a direct source, labeled as a confirmed platform requirement, confirmed legal requirement, evidence-backed tactic, or experimental practice. Priority 5 is critical. Priority 1 is experimental or optional. No item guarantees ranking, visibility, traffic, or citation.
GEO STRATEGY & GUARDRAILS
Is there a current crawler-to-purpose inventory?
Priority: 5
- What
- Confirmed platform requirement: Classify every platform behavior as automatic crawl, training crawl, user-triggered fetch, or answer-time reuse or grounding of indexed content.
- How
- Record each product token, user-agent pattern, control mode, owner, intended paths, and official source. Do not assume one rule governs the other modes.
- Why
- The same vendor can expose independent controls for access, training, user actions, and answer context.
- Reference
- OpenAI crawler overview
Has the organization approved separate retrieval and training policies?
Priority: 5
- What
- Evidence-backed tactic: Decide search visibility independently from foundation-model training.
- How
- Maintain a signed policy matrix for Googlebot/Google-Extended, OAI-SearchBot/GPTBot, Claude-SearchBot/ClaudeBot, and Applebot/Applebot-Extended.
- Why
- Blocking training need not sacrifice search.
- Reference
- Google common crawlers
Does every crawler-policy change have an accountable owner and review date?
Priority: 5
- What
- Evidence-backed tactic: Assign legal, security, SEO, and engineering responsibility.
- How
- Store intent, approver, effective date, and next review beside the rule.
- Why
- Fast-changing bot names and uses make undocumented rules unsafe.
- Reference
- Anthropic crawler controls
Are robots.txt rules version-controlled and tested before release?
Priority: 5
- What
- Evidence-backed tactic: Treat crawler policy as production configuration.
- How
- Diff parsed groups, test representative URLs per token, and retain rollback evidence.
- Why
- A malformed or overbroad rule can remove an entire answer surface.
- Reference
- Google robots.txt specification
Is the policy deployed on every relevant hostname and subdomain?
Priority: 5
- What
- Confirmed platform requirement: Cover each separately hosted locale, media host, CDN, docs site, and subdomain.
- How
- Fetch `/robots.txt` and target URLs on each host.
- Why
- Robots rules are host-scoped.
- Reference
- Anthropic crawler controls
Are automatic crawlers distinguished from user-triggered fetchers?
Priority: 5
- What
- Evidence-backed tactic: Document whether user fetches obey the same robots policy.
- How
- Test and record ChatGPT-User, Claude-User, Perplexity-User, and platform-specific semantics.
- Why
- A search-crawl allow or block may not govern a user request.
- Reference
- OpenAI crawler overview
Are official IP-range endpoints treated as changing dependencies?
Priority: 5
- What
- Evidence-backed tactic: Refresh allowlists rather than copying static IPs into policy forever.
- How
- Pull vendor JSON endpoints on a controlled schedule and alert on changes.
- Why
- Stale allowlists silently block legitimate bots.
- Reference
- Perplexity crawler documentation
Do stakeholder claims avoid promising inclusion, rank, or citation?
Priority: 5
- What
- Confirmed platform requirement: Describe controls as eligibility and discovery measures.
- How
- Remove guaranteed-visibility language from briefs and dashboards.
- Why
- Platforms reserve selection decisions.
- Reference
- Google AI features and websites
Is there a documented removal and escalation path per answer engine?
Priority: 4
- What
- Evidence-backed tactic: Cover emergency de-indexing, stale citations, copyright, and crawler malfunction.
- How
- Record platform tools, evidence required, contacts, and validation steps.
- Why
- Robots changes may not remove already indexed material immediately.
- Reference
- Anthropic blocking and removal
Are third-party search-index dependencies included in incident analysis?
Priority: 4
- What
- Evidence-backed tactic: Recognize that an answer engine may use partner indexes.
- How
- Check both the answer engine and underlying search-engine access/index state.
- Why
- Allowing one proprietary bot may not restore partner-fed results.
- Reference
- ChatGPT Search help
Does a recurring review compare English and Turkish visibility, accuracy, citations, referrals, experiments, incidents, and unresolved risks?
Priority: 5
- What
- Evidence-backed tactic: Give owners one balanced scorecard without collapsing unlike metrics.
- How
- Review platform-specific panels, uncertainty, language gaps, incidents, and next decisions with local owners.
- Why
- A complete GEO program joins measurement, operations, business outcomes, and governance across both audiences.
- Reference
- NIST AI 600-1 Generative AI Profile
AUDIENCE, INTENT & QUERY RESEARCH
Is every priority content asset tied to a researched user need?
Priority: 5
- What
- Evidence-backed tactic: Define the decision, task, or understanding a real audience needs.
- How
- Attach interview, support, sales, search, or product evidence to a concise need statement.
- Why
- Content built around an internal keyword list can miss the user's actual problem.
- Reference
- GOV.UK learning about user needs
Does the page have one explicit primary audience and purpose?
Priority: 5
- What
- Evidence-backed tactic: State whom the content serves and what successful use looks like.
- How
- Put audience, context, and intended outcome in the brief and test the finished page against them.
- Why
- A page that tries to serve incompatible audiences becomes ambiguous.
- Reference
- Google helpful content guidance
Are content opportunities grouped by information need rather than exact keyword wording?
Priority: 4
- What
- Evidence-backed tactic: Consolidate synonymous prompts that seek the same answer.
- How
- Map variants to one canonical need and split only when intent or required evidence differs.
- Why
- Google says AI systems understand synonyms and meanings, so variant-page proliferation adds little value.
- Reference
- Google AI optimization guide
Does the research cover informational, comparative, transactional, and action-oriented intents?
Priority: 4
- What
- Evidence-backed tactic: Represent the different jobs users ask an answer engine to perform.
- How
- Classify prompt evidence by intent and document the content or tool that resolves each class.
- Why
- A purely informational library misses decisions and next actions.
- Reference
- Google AI features and websites
Are likely follow-up questions mapped as a coherent journey?
Priority: 4
- What
- Evidence-backed tactic: Capture what a user needs before and after the initial answer.
- How
- Analyze support threads, conversational sessions, and related questions. Connect each stage with contextual links.
- Why
- Generative search is conversational and may issue additional targeted searches.
- Reference
- ChatGPT Search
Are complex questions decomposed into genuine subproblems?
Priority: 4
- What
- Evidence-backed tactic: Identify definitions, eligibility, options, evidence, risks, steps, and outcomes required for a complete answer.
- How
- Build a topic map and assign each subproblem to a section or linked specialist page.
- Why
- Google AI features can use query fan-out across subtopics and data sources.
- Reference
- Google AI features and websites
Does the content answer meaningful qualifiers and constraints?
Priority: 4
- What
- Evidence-backed tactic: Cover audience, budget, timing, geography, compatibility, eligibility, and risk qualifiers that change the answer.
- How
- Derive qualifiers from real decisions and show how recommendations change.
- Why
- ChatGPT may rewrite a prompt using location or remembered preferences.
- Reference
- ChatGPT Search
Are comparison questions organized around explicit decision criteria?
Priority: 4
- What
- Evidence-backed tactic: Name the dimensions that distinguish options.
- How
- Define audience, criteria, weights or tradeoffs, evidence date, and cases where each choice fits.
- Why
- Comparison content is more useful when readers can audit the reasoning.
- Reference
- Google helpful content guidance
Are “when not to,” exclusion, and no-answer needs represented?
Priority: 3
- What
- Evidence-backed tactic: Identify situations where the product, method, or recommendation does not apply.
- How
- Add contraindications, prerequisites, unsupported cases, and escalation paths.
- Why
- Completeness includes boundaries alongside persuasive positive claims.
- Reference
- Google helpful content guidance
Are actual grounding queries used to revise topic coverage?
Priority: 4
- What
- Evidence-backed tactic: Treat retrieval phrases reported by Microsoft as evidence of how systems map content.
- How
- Review Bing Webmaster Tools or Clarity grounding-query samples, validate user relevance, then fill material gaps.
- Why
- Grounding queries may differ from user wording.
- Reference
- Bing AI Performance
Do customer-facing teams contribute real questions and terminology?
Priority: 4
- What
- Evidence-backed tactic: Use language from support, sales, onboarding, research, and community interactions.
- How
- Maintain a consent-safe question repository with frequency, audience, resolution, and last-seen date.
- Why
- Real questions reveal needs and vocabulary that search tools miss.
- Reference
- GOV.UK learning about user needs
Are emerging questions distinguished from short-lived trend chasing?
Priority: 3
- What
- Evidence-backed tactic: Separate durable needs from news-driven spikes.
- How
- require audience fit, evidence availability, maintenance owner, and expected shelf life before publishing.
- Why
- Google warns against writing merely because a topic is trending.
- Reference
- Google helpful content guidance
Does each topic cluster have a clear canonical answer and specialist depth?
Priority: 3
- What
- Evidence-backed tactic: Prevent multiple pages from giving inconsistent partial answers.
- How
- name the canonical overview, link to detailed evidence, and retire or consolidate redundant pages.
- Why
- Coherent internal relationships help users traverse fan-out subtopics.
- Reference
- Google link best practices
Is a cross-engine, cross-language prompt panel maintained for priority needs?
Priority: 2
- What
- Experimental practice: Observe how named answer surfaces phrase and decompose the same need.
- How
- test natural prompts, several paraphrases, languages, locations, and dates. Retain outputs and null results.
- Why
- engines and contexts retrieve different source ecosystems.
- Reference
- Critical GEO survey
Does each answer make the user's safe next step explicit?
Priority: 4
- What
- Evidence-backed tactic: Close the information gap with a suitable action, decision aid, tool, source, or escalation.
- How
- test that the next step follows from the evidence and works without hidden prerequisites.
- Why
- People-first content should leave users feeling they learned enough to achieve their goal.
- Reference
- Google helpful content guidance
ANSWER ARCHITECTURE
Does the title accurately summarize the page's distinctive answer?
Priority: 5
- What
- Evidence-backed tactic: Use a descriptive, non-sensational title tied to the real content.
- How
- compare the title with the primary need, scope, and conclusion. Remove unsupported superlatives.
- Why
- Users and retrieval systems need a reliable statement of page purpose.
- Reference
- Google helpful content guidance
Is the primary answer or conclusion available before extended detail?
Priority: 4
- What
- Evidence-backed tactic: Give a concise, qualified answer early.
- How
- state the result, audience, material caveat, and evidence date, then expand.
- Why
- Readers can orient quickly while deeper context remains available.
- Reference
- Microsoft Bing AI Performance
Do section headings describe their actual topic or purpose?
Priority: 5
- What
- Evidence-backed tactic: Make headings predictive rather than clever or generic.
- How
- read the heading outline alone and verify it communicates the page's logic.
- Why
- Descriptive headings help people navigate and expose content relationships.
- Reference
- WCAG 2.2 headings and labels
Is the visible hierarchy encoded with real semantic headings, lists, and tables?
Priority: 5
- What
- Evidence-backed tactic: Preserve relationships in markup rather than relying on styling alone.
- How
- inspect the DOM and accessibility tree for heading levels, list semantics, table headers, and regions.
- Why
- Programmatic relationships survive alternate presentations and agent parsing.
Does each section resolve one coherent subquestion without arbitrary micro-fragmentation?
Priority: 4
- What
- Evidence-backed tactic: Use sections where topic boundaries naturally change.
- How
- merge tiny fragments that lack context and split only when a reader gains navigation or comparison value.
- Why
- Google explicitly says tiny content “chunks” are not required.
- Reference
- Google AI optimization guide
Can important paragraphs be understood with their necessary qualifiers intact?
Priority: 4
- What
- Evidence-backed tactic: Keep the claim, subject, condition, date, and source close together.
- How
- replace orphaned pronouns and detached caveats where ambiguity results.
- Why
- Reuse is safer when context is not stranded elsewhere.
- Reference
- Microsoft Bing AI Performance
Are central terms defined at first meaningful use?
Priority: 4
- What
- Evidence-backed tactic: State what a specialized term means in this context.
- How
- use a concise definition, expand acronyms, and link to a maintained glossary when needed.
- Why
- Explicit definitions reduce entity and concept ambiguity.
- Reference
- WHATWG `dfn` element
Are procedures presented as ordered, testable steps?
Priority: 4
- What
- Evidence-backed tactic: Distinguish prerequisites, actions, expected results, and failure paths.
- How
- use an ordered list with commands or UI labels that match the current product.
- Why
- Complete steps are easier to follow and verify than narrative fragments.
Do comparison tables have explicit dimensions, units, and accessible headers?
Priority: 4
- What
- Evidence-backed tactic: Encode row/column relationships and the basis of comparison.
- How
- use table captions, `th`, `scope` or equivalent associations. Note date and missing data.
- Why
- Microsoft recommends tables, while WCAG requires relationships to be programmatically determinable.
Are recommendations paired with tradeoffs and selection conditions?
Priority: 4
- What
- Evidence-backed tactic: Explain benefits, costs, risks, and who should choose differently.
- How
- add a decision rule and evidence behind each material recommendation.
- Why
- Unsupported universal advice is easy to quote incorrectly.
- Reference
- Google helpful content guidance
Are prerequisites and dependencies visible before the recommendation or procedure?
Priority: 4
- What
- Evidence-backed tactic: State required access, data, skills, versions, jurisdiction, and budget.
- How
- add a prerequisites block and link each dependency to authoritative instructions.
- Why
- Missing preconditions make an otherwise correct answer unusable.
- Reference
- Google helpful content guidance
Are limitations, edge cases, and unsupported scenarios explicit?
Priority: 4
- What
- Evidence-backed tactic: Bound the claim or method.
- How
- list known failure modes, exceptions, uncertainty, and escalation criteria beside the affected guidance.
- Why
- Caveats separated into a distant disclaimer are easily lost.
- Reference
- Google helpful content guidance
Are FAQ sections limited to real, distinct questions?
Priority: 3
- What
- Evidence-backed tactic: Include questions evidenced by users that add material information.
- How
- remove keyword-variant questions and merge answers that resolve the same need.
- Why
- Microsoft recommends FAQ sections, but Google rejects variant content written just for AI.
- Reference
- Google AI optimization guide
Are acronyms, symbols, and jargon expanded for the intended audience?
Priority: 3
- What
- Evidence-backed tactic: Make specialized language interpretable without erasing necessary precision.
- How
- define at first use and keep a controlled terminology list for repeated concepts.
- Why
- Clear language helps people process information quickly.
- Reference
- GOV.UK accessible documents
Are visible publication, review, and material-update signals unambiguous?
Priority: 4
- What
- Evidence-backed tactic: Distinguish when the page was published, reviewed, and substantively changed.
- How
- label dates, explain major changes, and align visible dates with structured values.
- Why
- Users need to judge currency without false freshness.
- Reference
- Google byline date guidance
EVIDENCE & CITATION INTEGRITY
Has the page's set of externally verifiable claims been inventoried?
Priority: 5
- What
- Evidence-backed tactic: Identify facts, numbers, comparisons, legal statements, and causal assertions needing support.
- How
- run a claim-level editorial pass and assign a source, owner, and review date.
- Why
- Page-level bibliographies cannot reveal unsupported individual claims.
- Reference
- Google fact-check guidance
Does each material claim use the strongest available primary source?
Priority: 5
- What
- Evidence-backed tactic: Prefer laws, standards, original research, official datasets, vendor documentation, or first-hand records.
- How
- trace secondary summaries back to the originating evidence and cite that page.
- Why
- Primary sources reduce distortion and make verification easier.
- Reference
- Google helpful content guidance
Do citations link to the deepest stable page that supports the claim?
Priority: 4
- What
- Evidence-backed tactic: Avoid home pages, search results, and generic documentation indexes.
- How
- link the exact report, section, dataset, specification, or release note. Preserve DOI or stable identifier where available.
- Why
- Direct sources reduce verification cost.
- Reference
- OpenAI web-search citations
Has a reviewer confirmed that every cited source actually supports the adjacent claim?
Priority: 5
- What
- Evidence-backed tactic: Test entailment, scope, and context rather than link presence.
- How
- open the source, locate the supporting passage, and record mismatches or missing qualifiers.
- Why
- AI systems can cite a page that does not support the generated statement.
- Reference
- Critical GEO survey
Are answer sentences decomposed into independently verifiable claims before citation scoring?
Priority: 5
- What
- Evidence-backed tactic: Identify each factual proposition and its associated citations.
- How
- Use a documented claim-splitting rule and human review for complex sentences.
- Why
- A 2023 snapshot of four generative search engines found only 51.5% of generated sentences were fully supported. Treat this as historical evidence of the risk rather than a current platform benchmark.
Are all externally verifiable answer claims checked for citation support?
Priority: 5
- What
- Evidence-backed tactic: Measure the share of claims fully supported by cited evidence.
- How
- Review each claim against its cited sources and record partial support separately.
- Why
- Fluent answers can contain uncited or only partly supported factual claims.
Does every citation actually support the claim it is attached to?
Priority: 5
- What
- Evidence-backed tactic: Measure the share of citations that entail or substantiate the associated claim.
- How
- Open the cited source, inspect the relevant passage, and score full, partial, or no support.
- Why
- In the same historical four-engine snapshot, 74.5% of citations fully supported the associated sentence. A real citation can still be irrelevant or misused.
Are citations placed next to the claims they support?
Priority: 4
- What
- Evidence-backed tactic: Minimize ambiguity between evidence and assertion.
- How
- add inline links or markers at sentence or paragraph level and keep a readable source list for full metadata.
- Why
- Claim-level proximity makes both human and automated attribution easier to audit.
- Reference
- OpenAI web-search citations
Are source publication date, version, and access date recorded where currency matters?
Priority: 5
- What
- Evidence-backed tactic: Make the evidence state reproducible.
- How
- show source date/version in notes and record the audit access date internally.
- Why
- Vendor behavior, law, prices, and research versions change.
- Reference
- Google byline date guidance
Does every statistic state its denominator, unit, period, and geography?
Priority: 5
- What
- Evidence-backed tactic: Preserve the minimum context needed to interpret a number.
- How
- attach sample/base, numerator and denominator, units, dates, location, and source beside the statistic.
- Why
- Context-free numbers are easy to reuse misleadingly.
- Reference
- Foundational GEO paper
Are study method, sample, population, and exclusions summarized with research claims?
Priority: 5
- What
- Evidence-backed tactic: Expose how the result was produced and whom it applies to.
- How
- add a method note and link to the full protocol or paper.
- Why
- A headline effect size without design context encourages false generalization.
- Reference
- Critical GEO survey
Do comparisons use the same baseline, unit, and observation window?
Priority: 4
- What
- Evidence-backed tactic: Make unlike measures visibly incomparable.
- How
- normalize where legitimate. Otherwise, label basis differences and avoid a synthetic rank.
- Why
- Comparison tables can manufacture certainty from mismatched inputs.
- Reference
- Google Dataset structured data
Are correlation, prediction, and causation distinguished?
Priority: 5
- What
- Evidence-backed tactic: Match claim language to the research design.
- How
- reserve causal verbs for randomized or defensible quasi-experimental evidence and explain confounding elsewhere.
- Why
- AI summaries can amplify an overclaimed causal sentence.
- Reference
- Critical GEO survey
Are uncertainty, ranges, and material disagreement reported?
Priority: 4
- What
- Evidence-backed tactic: Avoid presenting estimates or contested conclusions as settled facts.
- How
- include confidence intervals, plausible ranges, dissenting evidence, and what would change the conclusion.
- Why
- Factual reliability includes uncertainty.
- Reference
- Google helpful content guidance
Are quotations exact, bounded, and attributed to their original speaker and work?
Priority: 4
- What
- Evidence-backed tactic: Preserve wording without laundering a paraphrase as a quote.
- How
- verify the source, mark omissions or additions, link it, and name speaker, work, and date.
- Why
- The foundational GEO study's quotation result does not justify invented or decontextualized quotes.
- Reference
- WHATWG quotation semantics
Does the page represent credible conflicting sources rather than cherry-pick?
Priority: 4
- What
- Evidence-backed tactic: Surface material evidence that would change a reader's decision.
- How
- document source-selection criteria, summarize the disagreement, and explain the chosen conclusion.
- Why
- One-sided sourcing weakens trust and portability.
- Reference
- Google helpful content guidance
Are derived calculations reproducible?
Priority: 4
- What
- Evidence-backed tactic: Show formula, inputs, rounding, assumptions, and source data.
- How
- publish a worked example or downloadable notebook/spreadsheet for material calculations.
- Why
- Readers can verify the result and update it when inputs change.
- Reference
- Google Dataset structured data
Are original datasets available in documented, usable formats when disclosure is safe?
Priority: 4
- What
- Evidence-backed tactic: Publish the evidence behind original quantitative claims.
- How
- provide stable downloads or API access with schema, version, creator, license, coverage, and contact.
- Why
- Original research becomes verifiable and independently reusable.
- Reference
- Google Dataset structured data
Is dataset provenance explicit when data are copied, transformed, or aggregated?
Priority: 4
- What
- Evidence-backed tactic: Distinguish republication from derivation.
- How
- identify the canonical original with `sameAs` for an unchanged copy and `isBasedOn` for significant transformation or aggregation.
- Why
- Users need to trace data lineage.
- Reference
- Google Dataset provenance guidance
Are source licenses and reuse rights recorded?
Priority: 4
- What
- Evidence-backed tactic: Make legal reuse and attribution conditions visible for data, images, quotes, and code.
- How
- link the license, name creator/rightsholder, and retain required notices.
- Why
- Citation does not itself grant permission to reuse.
- Reference
- Google image license metadata
Are broken, redirected, retracted, and materially changed citations monitored?
Priority: 4
- What
- Evidence-backed tactic: Keep the evidence graph valid over time.
- How
- run link checks, preserve DOI/archive identifiers, review redirects, and flag retractions or changed conclusions.
- Why
- A live URL can still point to evidence that no longer supports the claim.
- Reference
- Schema.org `citation`
Are volatile claims assigned a review cadence and owner?
Priority: 5
- What
- Evidence-backed tactic: Identify prices, availability, policies, laws, product behavior, and “current” comparisons.
- How
- set source-specific expiry dates, alerts, and an update-or-remove workflow.
- Why
- Microsoft says accurate, up-to-date content matters for AI inclusion and citation.
- Reference
- Bing AI Performance
Is there a visible correction mechanism and material-change history?
Priority: 4
- What
- Evidence-backed tactic: Let readers report errors and see significant corrections.
- How
- publish contact/reporting paths, correction date, original error, corrected statement, and affected evidence.
- Why
- Traceable corrections improve accountability.
- Reference
- Google fact-check guidance
ORIGINALITY & INFORMATION GAIN
Does the page contribute a genuinely distinctive viewpoint?
Priority: 5
- What
- Evidence-backed tactic: Add analysis or experience unavailable in commodity summaries.
- How
- identify the page's unique claim, evidence, or framework in the brief and verify it survives editing.
- Why
- Google explicitly recommends a unique point of view for generative AI search.
- Reference
- Google AI optimization guide
Is claimed first-hand experience demonstrated with verifiable detail?
Priority: 5
- What
- Evidence-backed tactic: Show what was used, visited, tested, observed, or implemented.
- How
- include dates, setup, constraints, artifacts, and outcomes while protecting sensitive data.
- Why
- Google contrasts first-hand review with a restatement of existing content.
- Reference
- Google AI optimization guide
Does original research publish its question, method, data basis, and limitations?
Priority: 5
- What
- Evidence-backed tactic: Make proprietary findings inspectable.
- How
- link a methodology page and expose enough data or aggregates for verification.
- Why
- Original reporting is more valuable when reproducible and bounded.
- Reference
- Google helpful content guidance
Does sourced content add substantial analysis rather than paraphrase?
Priority: 5
- What
- Evidence-backed tactic: Go beyond rewriting another page.
- How
- compare the draft with cited sources and require a new synthesis, application, critique, dataset, or example.
- Why
- Google explicitly warns against copying or rewriting without added value.
- Reference
- Google helpful content guidance
Does the page provide material value beyond existing strong sources?
Priority: 4
- What
- Evidence-backed tactic: Identify the unresolved gap it closes.
- How
- benchmark leading primary and high-quality secondary sources, then document the incremental contribution.
- Why
- Commodity pages add little reason to retrieve or cite another source.
- Reference
- Google AI optimization guide
Do case studies disclose starting state, intervention, time window, and outcome?
Priority: 4
- What
- Evidence-backed tactic: Turn anecdotes into bounded evidence.
- How
- report baseline, context, steps, measurement, result, confounders, and what may not generalize.
- Why
- Decontextualized success stories invite false causal inference.
- Reference
- Google helpful content guidance
Are failures, negative results, and tradeoffs retained?
Priority: 4
- What
- Evidence-backed tactic: Publish what did not work where it changes the decision.
- How
- record failed approaches, conditions, costs, and lessons alongside successes.
- Why
- Negative evidence is distinctive and reduces survivorship bias.
- Reference
- Critical GEO survey
Can product tests or experiments be reproduced at the stated date and version?
Priority: 4
- What
- Evidence-backed tactic: Preserve setup, inputs, environment, and expected result.
- How
- publish test protocol, screenshots or logs, version numbers, and last rerun date.
- Why
- Product behavior drifts and unsupported “current” tests age quickly.
Are proprietary data claims accompanied by safe evidence artifacts?
Priority: 4
- What
- Evidence-backed tactic: Make internally derived insights auditable without leaking protected data.
- How
- publish aggregates, sampling rules, anonymization, schema, and a contact for access questions.
- Why
- Unsupported proprietary numbers are difficult to trust or reuse.
- Reference
- Google Dataset structured data
Are expert interviews linked to the interviewee's identity and original context?
Priority: 4
- What
- Evidence-backed tactic: Preserve who said what, when, and in what capacity.
- How
- obtain consent, publish transcript or relevant excerpt, link a profile, and separate opinion from fact.
- Why
- Attributable primary testimony is stronger than anonymous authority language.
- Reference
- Google helpful content guidance
Does the page use original examples that actually exercise the guidance?
Priority: 4
- What
- Evidence-backed tactic: Demonstrate the concept in a realistic scenario.
- How
- create examples from tested work, state assumptions, and verify all outputs.
- Why
- Original examples add practical information beyond definitions.
- Reference
- Google AI optimization guide
Does the page expose a reusable decision framework rather than a flat tips list?
Priority: 4
- What
- Evidence-backed tactic: Connect evidence, criteria, options, and outcomes.
- How
- publish a flow, rubric, calculator, or worked decision with limitations.
- Why
- A framework helps users apply expertise to their own context.
- Reference
- Google AI optimization guide
Is scaled, low-value derivative production prohibited?
Priority: 5
- What
- Confirmed platform requirement: Prevent mass-generated pages that add no value.
- How
- audit templates, automation, outsourcing, and network publishing for originality and human review.
- Why
- Google classifies scaled unoriginal content created to manipulate rankings as spam.
- Reference
- Google spam policies
Are update dates changed only after substantive review or modification?
Priority: 5
- What
- Evidence-backed tactic: Reject simulated freshness.
- How
- define what qualifies as a material update and retain review/change evidence.
- Why
- Google warns against changing dates merely to make pages appear fresh.
- Reference
- Google helpful content guidance
Is the reason for publishing primarily to help the intended audience?
Priority: 5
- What
- Evidence-backed tactic: Check that the page serves a real need beyond anticipated AI traffic.
- How
- require a user outcome, evidence gap, owner, and maintenance plan in the brief.
- Why
- Google asks whether content has a people-first purpose.
- Reference
- Google helpful content guidance
AUTHORSHIP & EDITORIAL TRUST
Does each expert or editorial page show a truthful visible byline?
Priority: 5
- What
- Evidence-backed tactic: Identify the accountable creator rather than a vague content team when a person is responsible.
- How
- place the byline near the title and link it to a maintained profile.
- Why
- Google asks who created content and recommends accurate authorship.
- Reference
- Google helpful content guidance
Does each author profile explain relevant experience and role?
Priority: 4
- What
- Evidence-backed tactic: Show why the creator is qualified for this topic.
- How
- list verifiable credentials, first-hand experience, affiliations, disclosures, selected work, and contact path.
- Why
- Expertise should be demonstrable rather than implied by tone.
- Reference
- Google ProfilePage structured data
Is the author's expertise matched to the page's subject and risk?
Priority: 5
- What
- Evidence-backed tactic: Avoid using a generic expert label across unrelated domains.
- How
- define topic scopes and require appropriate creator or reviewer qualifications for each risk class.
- Why
- Google asks whether an expert or enthusiast demonstrably knows the topic.
- Reference
- Google helpful content guidance
Are expert reviewers named with their exact review scope and date?
Priority: 5
- What
- Evidence-backed tactic: Distinguish writing, fact-checking, medical/legal review, and editorial approval.
- How
- display reviewer identity, qualification, reviewed sections, review date, and change trigger.
- Why
- A decorative “expert reviewed” badge can mislead.
- Reference
- Google helpful content guidance
Does every byline resolve to one stable canonical author URL?
Priority: 4
- What
- Evidence-backed tactic: Give each creator a durable identity page.
- How
- use the same profile URL in visible links and Article `author.url`. Redirect retired URLs.
- Why
- Google says `author.url` can uniquely identify the author.
- Reference
- Google Article structured data
Are all visible authors represented separately in Article markup?
Priority: 4
- What
- Confirmed platform requirement: Preserve multi-author accountability.
- How
- create one Person or Organization object per author and do not merge names into one field.
- Why
- Google explicitly instructs publishers to include every author in markup.
- Reference
- Google Article structured data
Does the site clearly identify its publication and publisher?
Priority: 5
- What
- Evidence-backed tactic: Explain the site's mission, editorial remit, publisher, and relationship to the organization.
- How
- maintain an About page and link it from content and navigation.
- Why
- Google cites site and publisher background as a trust cue.
- Reference
- Google helpful content guidance
For news content, are ownership, company, or network relationships transparent?
Priority: 5
- What
- Confirmed platform requirement: Identify the entity behind the publication.
- How
- disclose parent company, controlling organization, and material network relationships on an accessible page.
- Why
- Google News transparency policy asks for this information.
- Reference
- Google News policies
For news content, is usable contact information available?
Priority: 5
- What
- Confirmed platform requirement: Give readers a route to the publisher.
- How
- provide monitored editorial and correction contacts and, where appropriate, organization contact details.
- Why
- Google News transparency policy calls for contact information.
- Reference
- Google News policies
Is the editorial and fact-checking process published?
Priority: 4
- What
- Evidence-backed tactic: Explain source standards, review levels, update rules, and how independence is protected.
- How
- publish a concise policy and link it from high-stakes content.
- Why
- Readers can assess how claims reached publication.
- Reference
- Google fact-check guidance
Is there a public corrections policy with a working reporting path?
Priority: 5
- What
- Evidence-backed tactic: Define what gets corrected, how quickly, and how changes are shown.
- How
- test the reporting channel and sample recent corrections.
- Why
- Google requires a corrections policy or error-reporting mechanism for fact-check eligibility.
- Reference
- Google fact-check guidance
Are sponsorship, payment, affiliate interest, and material support clearly disclosed?
Priority: 5
- What
- Confirmed platform requirement: Separate commercial influence from independent editorial content.
- How
- label sponsorship near the affected content and explain the relationship.
- Why
- Google News forbids concealed or misrepresented sponsored content.
- Reference
- Google News policies
Is material automation or generative-AI use disclosed when readers would reasonably expect it?
Priority: 4
- What
- Evidence-backed tactic: Explain where automation contributed and what human checks occurred.
- How
- add a creation-method note covering generation, data, review, and limitations.
- Why
- Google recommends context about how automatically generated content was created.
Are conflicts of interest and relevant affiliations disclosed at claim level?
Priority: 4
- What
- Evidence-backed tactic: Reveal relationships that could affect interpretation.
- How
- collect author/reviewer disclosures and place material conflicts near the content.
- Why
- Transparent incentives help readers evaluate recommendations.
- Reference
- Google News policies
Does high-stakes advice use qualified review and authoritative local sources?
Priority: 5
- What
- Evidence-backed tactic: Apply stronger evidence and review to health, legal, financial, safety, and civic guidance.
- How
- require jurisdiction-appropriate primary sources, specialist review, update cadence, and escalation language.
- Why
- Error cost is high and rules change.
- Reference
- Google helpful content guidance
ENTITY IDENTITY & CONSISTENCY
Is there a governed registry for priority people, organizations, products, places, and concepts?
Priority: 5
- What
- Evidence-backed tactic: Maintain one internal record per entity.
- How
- assign a stable ID, canonical name, type, URL, aliases, relationships, source, owner, and review date.
- Why
- A registry prevents contradictory identities across pages and feeds.
- Reference
- Google Organization structured data
Is each entity's canonical public name consistent across the site?
Priority: 5
- What
- Evidence-backed tactic: Use the same primary identity in visible text, metadata, structured data, profiles, and feeds.
- How
- compare templates and data sources against the entity registry.
- Why
- Google recommends Organization names consistent with the site name.
- Reference
- Google Organization structured data
Are legitimate aliases, handles, abbreviations, and former names recorded?
Priority: 4
- What
- Evidence-backed tactic: Help readers reconcile names without treating every spelling as a separate entity.
- How
- show relevant aliases in profiles and use `alternateName` where supported.
- Why
- Google exposes `alternateName` for profile and organization identity.
- Reference
- Google ProfilePage structured data
Are ambiguous names explicitly disambiguated?
Priority: 5
- What
- Evidence-backed tactic: Distinguish entities sharing a name or acronym.
- How
- add type, location, affiliation, role, version, and canonical link at first mention.
- Why
- Names alone may not uniquely identify a person, company, product, or concept.
- Reference
- Google Organization structured data
Does each important entity have one stable canonical URL?
Priority: 5
- What
- Evidence-backed tactic: Give the entity a durable web identity.
- How
- choose the authoritative page, keep it current, link to it consistently, and redirect retired URLs.
- Why
- Google says an Organization URL helps uniquely identify it.
- Reference
- Google Organization structured data
Does the organization's canonical page state accurate identity and operating details?
Priority: 5
- What
- Evidence-backed tactic: Cover name, description, URL, logo, contact, address, legal identity, and relevant external profiles.
- How
- maintain these facts on the home or About page and reconcile them with markup.
- Why
- Google uses Organization data to understand administrative details and disambiguation.
- Reference
- Google Organization structured data
Does each marked-up profile focus on one affiliated Person or Organization?
Priority: 4
- What
- Confirmed platform requirement: Keep the profile's primary entity unambiguous.
- How
- use `ProfilePage.mainEntity` for that one person or organization and show the same focus visibly.
- Why
- Google lists single-entity focus as a ProfilePage content guideline.
- Reference
- Google ProfilePage structured data
Do Article author objects link to their canonical profile URLs?
Priority: 4
- What
- Evidence-backed tactic: Connect each work to a stable creator identity.
- How
- set `author.url` to the internal author page and use ProfilePage markup there when appropriate.
- Why
- Google recommends author URLs to uniquely identify creators.
- Reference
- Google Article structured data
Are `sameAs` links limited to high-confidence identity matches?
Priority: 4
- What
- Evidence-backed tactic: Link only external pages that unambiguously represent the same entity.
- How
- verify ownership, name, URL, and current status. Remove stale or fan-created profiles.
- Why
- Because `sameAs` is an identity assertion, keep generic related links elsewhere.
- Reference
- Google Organization structured data
Are official identifiers included only when verified and correctly scoped?
Priority: 3
- What
- Evidence-backed tactic: Use internal IDs, LEI, VAT, tax, GS1, ROR, DOI, or other identifiers for the right entity type.
- How
- validate against the issuing registry and record jurisdiction and format.
- Why
- Identifiers can disambiguate entities across systems.
- Reference
- Google Organization structured data
Are parent, subsidiary, brand, product, founder, employee, and ownership relationships explicit?
Priority: 4
- What
- Evidence-backed tactic: Model how entities relate rather than relying on name proximity.
- How
- state relationships visibly and encode supported properties consistently.
- Why
- Clear relationships prevent a model or reader from conflating a brand, legal entity, and product.
- Reference
- Schema.org Organization
Are people, organizations, products, events, places, and concepts typed correctly?
Priority: 5
- What
- Evidence-backed tactic: Avoid one generic entity object for unlike things.
- How
- audit visible nouns and JSON-LD types against the entity registry and source evidence.
- Why
- Correct typing preserves meaning and property validity.
- Reference
- Google structured data introduction
Does markup distinguish what a work is about from entities it merely mentions?
Priority: 4
- What
- Evidence-backed tactic: Separate the primary subject from incidental references.
- How
- use `about` or `mainEntity` for subject matter and `mentions` only for secondary referenced entities.
- Why
- Schema.org defines these relationships differently.
- Reference
- Schema.org `about`
Are expertise-topic assertions narrow, verifiable, and non-promotional?
Priority: 2
- What
- Experimental practice: Describe what a Person or Organization knows about without claiming unsupported mastery.
- How
- use visible evidence and, if useful, `knowsAbout` links or terms tied to real work.
- Why
- Schema.org says the property suggests possible expertise but does not imply it.
- Reference
- Schema.org `knowsAbout`
Are entity facts consistent across text, schema, images, video, and downloadable data?
Priority: 5
- What
- Evidence-backed tactic: Prevent cross-format contradictions in names, specs, dates, pricing, and claims.
- How
- compare all renditions against the canonical entity record during release.
- Why
- Microsoft recommends alignment across formats to reduce ambiguity.
- Reference
- Bing AI Performance
Are local business name, address, phone, hours, and service facts current everywhere?
Priority: 5
- What
- Evidence-backed tactic: Maintain one source of truth for local operating details.
- How
- reconcile website, markup, Bing Places, Google Business Profile, directories, and location pages.
- Why
- Microsoft highlights current address, hours, and contact information for location-based AI answers.
- Reference
- Bing AI Performance
Are current, legal, and former organization names distinguished while mergers, acquisitions, and rebrands are documented with dates?
Priority: 2
- What
- Experimental practice: Preserve identity continuity without conflating a current organization with former entities or names.
- How
- Keep `name`, `legalName`, `alternateName`, identifiers, and `sameAs` links consistent. Document dated changes visibly and update redirects and relationships.
- Why
- Consistent identity fields aid organization disambiguation. Treat dated corporate-history modeling as an operational inference because it is not a documented generative-search ranking factor.
Are product editions, models, plans, and versions unambiguously separated?
Priority: 5
- What
- Evidence-backed tactic: Prevent specifications or prices from crossing versions.
- How
- give each material variant a stable identifier, lifecycle dates, compatible versions, and canonical detail page.
- Why
- Entity ambiguity can produce wrong recommendations.
- Reference
- Google AI optimization guide
Do internal links name the destination entity or topic descriptively?
Priority: 4
- What
- Evidence-backed tactic: Replace generic anchors with concise, contextual identity labels.
- How
- make important entity pages reachable through real anchors whose text makes sense out of context.
- Why
- Google recommends descriptive, relevant anchor text.
- Reference
- Google link best practices
Are external identity sources periodically reconciled against first-party facts?
Priority: 4
- What
- Evidence-backed tactic: Find conflicts in trusted registries, profiles, reviews, and knowledge panels.
- How
- monitor high-impact external records, document the authoritative source, and correct owned surfaces or request fixes.
- Why
- AI answers can synthesize information from the brand site alongside other sources.
- Reference
- Google AI optimization guide
SEMANTIC HTML & STRUCTURED DATA
Are only documented applicable schema types used, without inventing a “GEO schema”?
Priority: 5
- What
- Confirmed platform requirement: Choose the most specific supported vocabulary for the real page.
- How
- validate against platform documentation and schema.org.
- Why
- Structured data supplies explicit clues but no special answer-engine type is documented.
- Reference
- Google structured data introduction
Is JSON-LD syntactically valid, reachable, and attached to the page it describes?
Priority: 4
- What
- Confirmed platform requirement: Use a maintainable supported format.
- How
- parse production markup and test required properties and URLs.
- Why
- Google recommends JSON-LD and requires access to the marked page.
- Reference
- Google structured data guidelines
Are required and useful recommended properties complete, accurate, and current?
Priority: 5
- What
- Confirmed platform requirement: Avoid partial or stale machine facts.
- How
- validate against the canonical database and expiry rules on every release.
- Why
- Incomplete required properties lose feature eligibility. Stale time-sensitive data may not display.
- Reference
- Google structured data guidelines
Are image and media URLs referenced by structured data crawlable and indexable?
Priority: 4
- What
- Confirmed platform requirement: Ensure machines can retrieve declared assets.
- How
- test status, robots, canonical, MIME, and stable URL from the edge.
- Why
- Google cannot use inaccessible referenced images.
- Reference
- Google structured data guidelines
Is paywalled content declared with accurate platform-supported markup rather than crawler-only cloaking?
Priority: 5
- What
- Confirmed platform requirement: Distinguish restricted access honestly.
- How
- implement Google `isAccessibleForFree`/`hasPart` where applicable and Apple's page-level signal.
- Why
- Correct markup preserves eligibility while communicating restrictions.
- Reference
- Google paywalled-content markup
Are visual relationships also programmatically determinable?
Priority: 5
- What
- Evidence-backed tactic: Encode headings, lists, definitions, tables, landmarks, and labels with appropriate HTML.
- How
- inspect DOM and accessibility-tree output, including content injected by components.
- Why
- Semantic relationships remain available when presentation changes.
Are citations and internal references real crawlable links?
Priority: 5
- What
- Evidence-backed tactic: Use `<a href>` rather than click handlers, styled spans, or script-only navigation.
- How
- inspect rendered HTML and test links without client JavaScript.
- Why
- Links expose a durable relationship and destination.
- Reference
- Google link best practices
Are self-contained charts, diagrams, code listings, and quotations grouped with `figure` and `figcaption`?
Priority: 3
- What
- Evidence-backed tactic: Bind an asset to its caption and stable label.
- How
- use semantic figure markup and refer to it by label rather than “above” or “below.”
- Why
- WHATWG defines figures as self-contained units with optional captions.
- Reference
- WHATWG figure semantics
Are block and inline quotations marked and attributed correctly?
Priority: 3
- What
- Evidence-backed tactic: Distinguish quoted material from author prose.
- How
- use `blockquote` or `q` where appropriate, a valid source URL when useful, and visible attribution outside the quote.
- Why
- HTML semantics preserve quotation boundaries.
- Reference
- WHATWG blockquote semantics
Is the `cite` element used for a work title rather than a person or quotation?
Priority: 3
- What
- Evidence-backed tactic: Keep citation semantics valid.
- How
- mark only the title of the referenced work with `cite` and provide a separate clickable source.
- Why
- WHATWG explicitly limits `cite` to titles of works.
- Reference
- WHATWG `cite` semantics
Are defining instances encoded and linkable where definitions matter?
Priority: 3
- What
- Evidence-backed tactic: Make a term and its definition explicitly related.
- How
- use `dfn`, a stable fragment ID, and contextual links from later uses when helpful.
- Why
- WHATWG defines `dfn` as the defining instance of a term.
- Reference
- WHATWG `dfn` semantics
Are publication, modification, event, and expiry dates machine-readable and correctly labeled?
Priority: 4
- What
- Evidence-backed tactic: Separate different temporal facts.
- How
- use visible labels, valid `<time datetime>` values, correct timezone where relevant, and matching structured data.
- Why
- The time element provides a machine-readable representation.
- Reference
- WHATWG `time` semantics
Does structured data describe content users can actually see?
Priority: 5
- What
- Confirmed platform requirement: Prevent hidden, misleading, or contradictory claims in JSON-LD.
- How
- compare every material property with the canonical rendered page and source data.
- Why
- Google requires markup to represent visible page content.
- Reference
- Google structured data guidelines
Is the most specific accurate, Google-supported page type used for the intended feature?
Priority: 4
- What
- Evidence-backed tactic: Avoid generic or invented schema where a documented type applies.
- How
- start from the relevant Google feature guide, then add broader Schema.org only for truthful secondary meaning.
- Why
- Google says Search documentation is definitive for Google behavior.
- Reference
- Google structured data introduction
Are fewer complete, accurate properties favored over broad incomplete markup?
Priority: 4
- What
- Evidence-backed tactic: Prioritize data quality.
- How
- populate required and meaningful recommended properties from governed sources. Omit unknown values.
- Why
- Google recommends supplying fewer complete and accurate properties rather than every possible property.
- Reference
- Google structured data introduction
Do Article author, date, headline, and image values match the article?
Priority: 4
- What
- Confirmed platform requirement: Keep creator and publication metadata aligned.
- How
- validate every author separately, stable author URLs, accurate dates, and representative crawlable images.
- Why
- Google's Article guide defines the supported fields and author best practices.
- Reference
- Google Article structured data
Is Organization markup placed on the canonical home or About page rather than repeated indiscriminately?
Priority: 3
- What
- Evidence-backed tactic: Centralize identity details.
- How
- maintain one complete Organization node and reference its stable `@id` from other markup.
- Why
- Google recommends organization details on a single home or About page.
- Reference
- Google Organization structured data
Are first-party datasets described with discoverable metadata?
Priority: 3
- What
- Evidence-backed tactic: Publish name, description, creator, identifier, license, coverage, distribution, and provenance.
- How
- add Dataset/DataDownload markup to the canonical dataset landing page and validate it.
- Why
- Google uses structured metadata for Dataset Search discovery.
- Reference
- Google Dataset structured data
Are `citation` and `isBasedOn` used only when they truthfully express work lineage?
Priority: 1
- What
- Experimental practice: Encode a reference or derivation in addition to visible citations.
- How
- add these Schema.org properties to CreativeWork objects and validate consumers before scaling.
- Why
- The vocabulary can make provenance relationships explicit.
- Reference
- Schema.org `citation`
Is structured data validated without being sold as an AI citation switch?
Priority: 5
- What
- Evidence-backed tactic: Test syntax, feature eligibility, visible consistency, and post-deploy errors.
- How
- use validator and platform reports, then measure outcomes.
- Why
- Google says structured data is not required for generative AI search and there is no special AI schema.
- Reference
- Google AI optimization guide
MULTIMODAL CONTENT & PROVENANCE
Do images and videos add evidence or explanation rather than decoration alone?
Priority: 4
- What
- Evidence-backed tactic: Use media that helps answer the user's need.
- How
- map each asset to a claim, procedure, comparison, or demonstration and remove redundant stock imagery.
- Why
- Google recommends high-quality relevant media for generative AI search opportunities.
- Reference
- Google AI optimization guide
Does every informative image have useful contextual alt text?
Priority: 5
- What
- Evidence-backed tactic: Convey the image's purpose and the information needed to understand it.
- How
- write concise alt text in page context. Use empty alt for decorative images and avoid keyword stuffing.
- Why
- Google calls alt text its most important image metadata and WCAG requires text alternatives.
- Reference
- Google Images best practices
Is each important image placed near explanatory text and a factual caption?
Priority: 4
- What
- Evidence-backed tactic: Connect media to the surrounding claim.
- How
- place it in the relevant section, name entities, date the asset when material, and cite its source.
- Why
- Google derives image subject matter from nearby page content, captions, and titles.
- Reference
- Google Images best practices
Do charts and diagrams have a complete text or data equivalent?
Priority: 5
- What
- Evidence-backed tactic: Expose trends, labels, units, and conclusions outside pixels.
- How
- provide concise alt plus nearby interpretation, a data table, or long description when alt is insufficient.
- Why
- WCAG says complex non-text content needs an alternative that presents equivalent information.
- Reference
- WCAG non-text content
Are figures given stable captions and referenced by label rather than position?
Priority: 3
- What
- Evidence-backed tactic: Preserve meaning across responsive layouts and extraction.
- How
- use `figure`, `figcaption`, an ID, and references such as “Figure 2.”
- Why
- WHATWG recommends labels instead of “above” or “below” references.
- Reference
- WHATWG figure semantics
Is the preferred preview image representative, specific, and high quality?
Priority: 4
- What
- Evidence-backed tactic: Avoid generic logos or misleading thumbnails.
- How
- select a relevant image, expose consistent image metadata, and review crops and aspect ratios.
- Why
- Google recommends representative, non-generic, high-resolution preferred images.
- Reference
- Google Images best practices
Are image filenames, alt text, captions, and embedded text localized?
Priority: 3
- What
- Evidence-backed tactic: Adapt the whole asset for its audience.
- How
- translate meaningful metadata and recreate text-heavy graphics rather than overlaying partial translations.
- Why
- Google recommends translating localized image filenames and contextual metadata.
- Reference
- Google Images best practices
Do important media assets use stable, crawlable canonical URLs?
Priority: 4
- What
- Evidence-backed tactic: Preserve discoverability and attribution across reuse.
- How
- avoid expiring signed URLs for public assets, redirect replacements, and use the same image URL when the same asset recurs.
- Why
- Google recommends consistent image URLs and stable video URLs.
- Reference
- Google Images best practices
Are creator, credit, copyright, license, and acquisition details attached to reusable assets?
Priority: 4
- What
- Evidence-backed tactic: Preserve attribution and rights.
- How
- add IPTC or structured metadata, a visible credit, license URL, and licensor route where relevant.
- Why
- Google can show creator and licensing details in Images.
- Reference
- Google image license metadata
Are AI-generated ecommerce images and product data labeled with the required metadata?
Priority: 5
- What
- Confirmed platform requirement: Identify synthetic product media and generated catalog fields.
- How
- embed IPTC `DigitalSourceType` `TrainedAlgorithmicMedia` for AI images and submit generated product fields separately as labeled AI content.
- Why
- Google Merchant Center policy requires this treatment.
Where the workflow supports it, do high-risk original media assets carry verifiable Content Credentials?
Priority: 2
- What
- Evidence-backed tactic: Add cryptographically verifiable, tamper-evident provenance for origin and edits.
- How
- Create a C2PA 2.4 manifest, bind and sign it, preserve validated ingredient provenance through authorized transformations, and validate the delivered asset.
- Why
- C2PA standardizes signed provenance assertions, content bindings, and validation states for media workflows.
Does each primary video have a dedicated, stable watch page when appropriate?
Priority: 3
- What
- Evidence-backed tactic: Give a video its own canonical context.
- How
- create a page where that video is the main content with a unique title and description.
- Why
- Google requires a dedicated watch page for eligibility in video features.
- Reference
- Google Video best practices
Do video metadata and markup accurately describe the actual video?
Priority: 4
- What
- Confirmed platform requirement: Align title, description, thumbnail, dates, content URL, and regions with the media.
- How
- validate VideoObject and compare every field with the playable asset.
- Why
- Google requires structured data to be consistent with video content and other metadata.
- Reference
- Google Video best practices
Do prerecorded audio and video provide accurate captions and descriptive transcripts?
Priority: 5
- What
- Evidence-backed tactic: Expose speech, relevant sounds, and visual information needed for understanding in text.
- How
- human-review captions, publish a transcript, identify speakers, and include descriptions needed to understand visuals.
- Why
- W3C identifies captions and transcripts as accessibility alternatives.
Are video chapters, timestamps, transcript claims, and surrounding page text aligned?
Priority: 4
- What
- Evidence-backed tactic: Make key moments navigable without cross-format contradiction.
- How
- add accurate chapter labels or Clip/SeekToAction data and reconcile dates, names, numbers, and conclusions across formats.
- Why
- Google supports key moments and Microsoft recommends cross-format entity consistency.
- Reference
- Google Video key moments
LOCALIZATION & MARKET CONTEXT
Are locale alternates correctly annotated and mutually coherent?
Priority: 5
- What
- Confirmed platform requirement: Connect language/region equivalents.
- How
- validate `hreflang` codes, self/return references, canonicals, and sitemap or HTML annotations.
- Why
- It helps Google select the right locale URL.
- Reference
- Localized page versions
Are robots, meta, status, canonical, and structured-data rules consistent across locale variants?
Priority: 5
- What
- Confirmed platform requirement: Prevent one language from silently losing eligibility.
- How
- run the full retrieval matrix for every locale.
- Why
- Google explicitly asks locale-adaptive sites to apply robots controls consistently.
- Reference
- Locale-adaptive crawling
Is important localized content available without IP or `Accept-Language` dependence and without forced redirects?
Priority: 5
- What
- Confirmed platform requirement: Serve stable locale URLs to any valid crawler.
- How
- test from US and target regions with missing/different language headers.
- Why
- Googlebot often uses US IPs and sends no `Accept-Language`.
- Reference
- Locale-adaptive crawling
Does each language version have a distinct stable URL?
Priority: 5
- What
- Evidence-backed tactic: Make localized content independently addressable.
- How
- use locale-specific paths or hosts and avoid cookie-only or header-only language swaps.
- Why
- Google recommends different URLs because dynamic variations may not all be crawled.
- Reference
- Google multilingual site guidance
Do all `hreflang` variants list themselves and every reciprocal version with absolute URLs?
Priority: 5
- What
- Confirmed platform requirement: Keep the locale cluster complete and bidirectional.
- How
- generate one consistent set, validate language-region codes, and test return links.
- Why
- Google may ignore non-reciprocal or incomplete annotations.
- Reference
- Google localized versions guidance
Is an `x-default` fallback declared for unmatched language or region users?
Priority: 4
- What
- Evidence-backed tactic: Identify the neutral selector or default page.
- How
- add `hreflang="x-default"` to the same alternate set and verify its purpose.
- Why
- Google documents `x-default` for users whose locale is not explicitly targeted.
- Reference
- Google localized versions guidance
Does each page use one clear primary visible language for content and navigation?
Priority: 5
- What
- Evidence-backed tactic: Avoid mixed-language pages and boilerplate-only translation.
- How
- audit main content, menus, widgets, errors, captions, and structured text.
- Why
- Google determines language from visible content and recommends a single language per page.
- Reference
- Google multilingual site guidance
Is localized content adapted by a fluent human for local intent and conventions?
Priority: 5
- What
- Evidence-backed tactic: Preserve meaning rather than mirror sentence structure mechanically.
- How
- review terminology, examples, tone, legal context, sources, and user tasks with a locale expert.
- Why
- A translated shell around unchanged content is a poor user experience.
- Reference
- Google multilingual site guidance
Can users switch language or region without forced automatic redirection?
Priority: 5
- What
- Evidence-backed tactic: Preserve user control and linkability.
- How
- provide visible alternate links, remember a preference without hiding URLs, and avoid IP- or language-forced redirects.
- Why
- Google warns automatic redirects can prevent users and crawlers from seeing variants.
- Reference
- Google multilingual site guidance
Is the document language declared with a valid `lang` value on the root element?
Priority: 5
- What
- Confirmed platform requirement: Expose the page's default human language.
- How
- use a valid BCP 47 tag on `<html>` and test generated layouts.
- Why
- W3C identifies the `lang` attribute as the standard declaration for text processing and accessibility.
- Reference
- W3C declaring language in HTML
Are inline language changes and bidirectional text marked correctly?
Priority: 4
- What
- Confirmed platform requirement: Preserve pronunciation, shaping, and reading order.
- How
- add `lang` to passages in another language and use appropriate `dir`, `bdi`, or `bdo` handling.
- Why
- W3C guidance requires declaration at the highest applicable level and direction separately.
- Reference
- W3C declaring language in HTML
Are currency, tax, units, dates, time zones, availability, and examples locally correct?
Priority: 5
- What
- Evidence-backed tactic: Adapt facts that change the answer by market.
- How
- source values per locale, display units and conversion basis, and date volatile details.
- Why
- Location is part of answer relevance and ChatGPT can use approximate or precise location.
- Reference
- ChatGPT Search
Do legal, regulatory, medical, and civic claims cite authoritative sources for the target jurisdiction?
Priority: 5
- What
- Evidence-backed tactic: Prevent one country's rule from being generalized globally.
- How
- name jurisdiction, effective date, issuing authority, and exceptions near the claim.
- Why
- Local context materially changes high-stakes answers.
- Reference
- Google multilingual site guidance
Are local address, phone, hours, service area, and contact methods consistent?
Priority: 5
- What
- Evidence-backed tactic: Keep location entities accurate in every locale.
- How
- reconcile page text, schema, platform listings, maps, and local profiles from one governed source.
- Why
- Google lists local addresses, phone numbers, currency, and local links among locale signals.
- Reference
- Google multilingual site guidance
Are local names, transliterations, grammatical forms, and aliases reconciled to the same entity?
Priority: 4
- What
- Evidence-backed tactic: Connect language-specific labels without erasing native usage.
- How
- maintain locale-aware canonical names and alternate names, and link to the same stable entity ID.
- Why
- Entity resolution can fail when translations look like separate entities.
- Reference
- Google ProfilePage structured data
Does each localized page cite sources credible and accessible to that audience?
Priority: 4
- What
- Evidence-backed tactic: Avoid importing all evidence from another language or market.
- How
- prefer authoritative local primary sources, retain global sources where applicable, and explain cross-market transfer.
- Why
- Search and AI source ecosystems vary by language and geography.
- Reference
- Cross-engine GEO study
Are priority prompts tested across language, country, city, and personalization contexts?
Priority: 2
- What
- Experimental practice: Detect where answers, citations, and entity matching diverge.
- How
- run named locales, natural local-language paraphrases, repeated dates, and controlled locations. Record null results.
- Why
- OpenAI documents location-aware rewriting and research finds cross-language variation.
- Reference
- ChatGPT Search
For same-language regional duplicates, are canonical and `hreflang` signals coordinated?
Priority: 5
- What
- Confirmed platform requirement: Choose a preferred duplicate while preserving regional targeting.
- How
- point each regional version to the intended canonical and maintain the complete alternate set.
- Why
- Google explicitly recommends canonical plus `hreflang` for same-language regional duplicates.
- Reference
- Google multilingual site guidance
GOOGLE AI SEARCH: VERIFIED 23 JULY 2026
Is each target page indexed and eligible to appear with a snippet in Google Search?
Priority: 5
- What
- Confirmed platform requirement: Establish baseline AI-feature eligibility.
- How
- Inspect canonical/index status and snippet controls in Search Console.
- Why
- This is Google's stated technical gate for supporting links.
- Reference
- Google AI features and websites
Can Googlebot crawl every target page and required resource?
Priority: 5
- What
- Confirmed platform requirement: Allow Search crawling where AI visibility is desired.
- How
- Test robots, HTTP response, rendered HTML, CSS, JS, and media.
- Why
- Googlebot is the access control for AI features inside Search.
- Reference
- Google AI features and websites
Is `nosnippet` absent where Google AI citation eligibility is desired?
Priority: 5
- What
- Confirmed platform requirement: Audit page and header directives.
- How
- Check HTML and `X-Robots-Tag` on canonical and duplicate responses.
- Why
- `nosnippet` prevents direct input to AI Overviews and AI Mode.
- Reference
- Google robots meta specifications
Is `max-snippet` intentionally sized for the desired AI-preview policy?
Priority: 4
- What
- Confirmed platform requirement: Avoid accidental zero or overly restrictive values.
- How
- Inventory effective directives after proxy/CDN composition.
- Why
- Google applies the limit to direct input for AI Overviews and AI Mode.
- Reference
- Google robots meta specifications
Is `data-nosnippet` applied only to content intentionally excluded from Google AI previews?
Priority: 4
- What
- Confirmed platform requirement: Protect selected text without suppressing the entire page.
- How
- Put the attribute on valid `div`, `span`, or `section` elements in initial/rendered DOM.
- Why
- It controls text-level preview use.
- Reference
- Google robots meta specifications
Is `noindex` absent from pages intended for Google AI features?
Priority: 5
- What
- Confirmed platform requirement: Check HTML and non-HTML headers.
- How
- Compare origin, CDN, mobile, and rendered responses.
- Why
- A page must remain eligible for Search indexing.
- Reference
- Google AI features and websites
Is Google-Extended treated separately from Google Search AI features?
Priority: 5
- What
- Confirmed platform requirement: Do not use it as an AI Overviews/AI Mode switch.
- How
- Set it only for Gemini Apps/Vertex grounding and future Gemini training policy.
- Why
- It does not govern Search inclusion.
- Reference
- Google common crawlers
Is GoogleOther excluded from Google Search visibility decisions?
Priority: 3
- What
- Confirmed platform requirement: Avoid calling it the AI Overviews crawler.
- How
- Classify it as generic product-team/R&D fetching unless a specific official use is documented.
- Why
- Its rules affect no specific product.
- Reference
- Google common crawlers
Are Google inspection-tool results distinguished from production Googlebot access?
Priority: 4
- What
- Confirmed platform requirement: Test both tooling and actual crawl/index evidence.
- How
- Avoid WAF rules that only permit `Google-InspectionTool`.
- Why
- That token affects tests but does not affect Search.
- Reference
- Google common crawlers
Is suspected Googlebot traffic verified beyond its user-agent string?
Priority: 5
- What
- Confirmed platform requirement: Defend against spoofing without blocking genuine crawls.
- How
- Use reverse DNS or Google's published ranges.
- Why
- User-agent strings are spoofable.
- Reference
- Googlebot documentation
Does the mobile response contain equivalent primary content, metadata, links, and structured data?
Priority: 5
- What
- Confirmed platform requirement: Audit what smartphone Googlebot sees.
- How
- Compare rendered mobile and desktop DOMs and response directives.
- Why
- Google primarily indexes mobile content.
- Reference
- Mobile-first indexing best practices
Are related subtopics independently crawlable for query fan-out?
Priority: 4
- What
- Evidence-backed tactic: Make supporting pages and sections discoverable through normal links.
- How
- Crawl the topic cluster from its hub and inspect canonicals/index state.
- Why
- AI Mode and AI Overviews may issue multiple related searches.
- Reference
- Google AI features and websites
Is Google generative-AI visibility measured in the correct Search Console views?
Priority: 4
- What
- Confirmed platform requirement: Use the dedicated report when available and retain the overall Web report.
- How
- Record rollout availability and export comparable date ranges.
- Why
- AI feature data remains included in overall performance.
Are URL, country, device, and time dimensions reviewed in Google's generative-AI report?
Priority: 3
- What
- Confirmed platform requirement: Separate visibility shifts by surface context.
- How
- Monitor pages, countries, devices, and hourly/daily/weekly/monthly trends available to the property.
- Why
- Aggregate totals can hide regional or template failures.
After preview-control changes, was recrawl and processing time allowed and verified?
Priority: 5
- What
- Confirmed platform requirement: Do not call a change ineffective immediately.
- How
- inspect fetched HTML, request recrawl, and monitor until the canonical is reprocessed.
- Why
- Google says changes can take days to months to recrawl.
- Reference
- Google AI features and websites
Has the team rejected any alleged special Google “GEO requirement”?
Priority: 5
- What
- Confirmed platform requirement: Keep Google AI eligibility grounded in Search technical requirements and policies.
- How
- Require an official source before adding a special tag, file, or markup.
- Why
- Google states there are no additional technical requirements or special optimizations.
- Reference
- Google AI features and websites
CHATGPT & OPENAI: VERIFIED 23 JULY 2026
Is OAI-SearchBot allowed on every page intended for ChatGPT Search summaries and citations?
Priority: 5
- What
- Confirmed platform requirement: Permit the search-specific bot.
- How
- Test the effective robots group for canonical content and assets.
- Why
- Opted-out sites are not shown in Search answers, though navigational links may remain.
- Reference
- OpenAI crawler overview
Does the CDN/WAF allow current OAI-SearchBot IP ranges as well as its user agent?
Priority: 5
- What
- Confirmed platform requirement: Remove network-layer false blocks.
- How
- consume `https://openai.com/searchbot.json`, combine IP and UA conditions, and test origin logs.
- Why
- OpenAI explicitly requires host/CDN access from published ranges.
- Reference
- ChatGPT Search help
Do user-agent rules tolerate OAI-SearchBot version changes and the robots marker?
Priority: 4
- What
- Confirmed platform requirement: Match the stable product token instead of a frozen full string.
- How
- test normal and `robots.txt`-marker request forms.
- Why
- OpenAI says version numbers may change and robots fetches may carry an extra marker.
- Reference
- OpenAI crawler overview
Is GPTBot policy independent from OAI-SearchBot policy?
Priority: 5
- What
- Confirmed platform requirement: Decide training use without accidentally suppressing Search.
- How
- give each token its own robots group and tests.
- Why
- OpenAI permits Search while disallowing model-training crawling.
- Reference
- OpenAI crawler overview
Is ChatGPT-User excluded from automatic Search inclusion logic?
Priority: 4
- What
- Confirmed platform requirement: Do not use it as the Search allow/deny token.
- How
- manage Search through OAI-SearchBot and treat ChatGPT-User as user-triggered access.
- Why
- OpenAI says ChatGPT-User does not determine Search appearance.
- Reference
- OpenAI crawler overview
Has the site tested user-triggered ChatGPT access separately from robots-controlled crawling?
Priority: 4
- What
- Confirmed platform requirement: Observe actions from ChatGPT-User.
- How
- test representative public URLs and inspect authenticated/WAF behavior.
- Why
- robots.txt may not apply to user-initiated actions.
- Reference
- OpenAI crawler overview
Was the documented robots-policy propagation window considered after an OpenAI change?
Priority: 4
- What
- Confirmed platform requirement: Avoid premature pass/fail judgments.
- How
- retest after at least the stated adjustment window and confirm live requests.
- Why
- OpenAI says Search systems may take about 24 hours to adjust.
- Reference
- OpenAI crawler overview
If even title/link exposure is prohibited, is `noindex` used on a crawlable page rather than relying only on OAI-SearchBot disallow?
Priority: 5
- What
- Confirmed platform requirement: Align the control with the removal goal.
- How
- permit the bot to read `noindex`, then verify de-indexing.
- Why
- OpenAI may surface a disallowed URL/title learned elsewhere.
- Reference
- OpenAI publisher FAQ
Is ChatGPT referral traffic identified by its documented campaign parameter?
Priority: 4
- What
- Confirmed platform requirement: Preserve and report `utm_source=chatgpt.com`.
- How
- prevent redirects from stripping it and segment landing/conversion analytics.
- Why
- OpenAI adds this parameter to referral URLs.
- Reference
- OpenAI publisher FAQ
Are intended ChatGPT Search pages publicly reachable without login, CAPTCHA, or consent dead ends?
Priority: 5
- What
- Confirmed platform requirement: Ensure a usable public response.
- How
- fetch through CDN/origin with the verified bot path and a clean session.
- Why
- OpenAI describes public websites as eligible and separately flags auth/challenge blocks.
- Reference
- OpenAI advertiser crawler guidance
Is ChatGPT Search tested against exact branded prompts as well as likely rewrites and follow-up queries?
Priority: 4
- What
- Evidence-backed tactic: Validate discovery through query variants.
- How
- maintain a prompt set covering subtopics, recency, comparison, and locale.
- Why
- ChatGPT may rewrite one prompt into multiple targeted partner queries.
- Reference
- ChatGPT Search help
For commerce sites, is current first-party product data supplied through supported integrations or feeds where eligible?
Priority: 4
- What
- Evidence-backed tactic: Reduce stale product facts.
- How
- validate Shopify Catalog integration or apply for direct product-feed access and reconcile feed-to-page data.
- Why
- OpenAI says direct feeds help reflect current products.
- Reference
- Shopping with ChatGPT Search
Do commerce fields such as availability, price, descriptions, and merchant identity agree across feed, structured metadata, and landing page?
Priority: 4
- What
- Confirmed platform requirement: Eliminate conflicting product facts.
- How
- diff catalog exports against canonical pages on every update.
- Why
- ChatGPT uses merchant and third-party metadata and ranks merchant options partly on availability and price.
- Reference
- Shopping with ChatGPT Search
Are interactive pages semantically operable for ChatGPT Agent in Atlas?
Priority: 3
- What
- Evidence-backed tactic: Make buttons, menus, forms, roles, labels, and states machine-readable.
- How
- audit against WAI-ARIA patterns and run representative agent journeys.
- Why
- OpenAI says Atlas uses ARIA semantics to interpret interfaces.
- Reference
- OpenAI publisher FAQ
BING & COPILOT: VERIFIED 23 JULY 2026
Can Bingbot crawl target content and resources?
Priority: 5
- What
- Confirmed platform requirement: Establish Bing index eligibility feeding Bing and AI-powered experiences.
- How
- test robots, live fetch, response, and rendered markup in Bing Webmaster Tools.
- Why
- Bing ties Copilot discovery to crawl/index availability.
- Reference
- Bing sitemaps in AI-powered search
If a specific Bingbot group exists, does it repeat needed generic directives?
Priority: 4
- What
- Confirmed platform requirement: Avoid unintended policy gaps.
- How
- parse the effective Bingbot group rather than assuming `*` merges into it.
- Why
- Bing says a specific group causes it to ignore generic directives.
- Reference
- Bing robots.txt guidance
Is Bingbot allowed to crawl pages carrying `noindex` until Bing can observe the directive?
Priority: 5
- What
- Confirmed platform requirement: Use the correct removal mechanism.
- How
- remove robots disallow, serve `noindex`, and monitor index state.
- Why
- Bing must fetch the page to read `noindex`.
- Reference
- Block URLs from Bing
Is Microsoft's generative-model training preference deliberately configured rather than inferred from crawl access?
Priority: 4
- What
- Confirmed platform requirement: Audit Bing-supported meta/X-Robots controls, including the currently documented training-use directive.
- How
- record the exact effective tag and verify in fetched HTML/header.
- Why
- Bing documents content-display and generative-AI data-use controls separately.
- Reference
- Bing robots meta tags
If IndexNow-participating engines matter in target markets, are new, meaningfully changed, and deleted URLs submitted?
Priority: 3
- What
- Evidence-backed tactic: Use IndexNow as an optional freshness notification alongside crawlable links and XML sitemaps.
- How
- Enable a trusted CMS or CDN integration, or deploy a verified API key, and submit only recent lifecycle changes.
- Why
- Participating engines can receive change notifications sooner, but each engine still decides whether and when to crawl or index a URL.
- Reference
- IndexNow protocol documentation
Are IndexNow and XML sitemaps used together rather than treated as substitutes?
Priority: 5
- What
- Evidence-backed tactic: Pair real-time change notification with comprehensive inventory.
- How
- compare submitted-change logs with sitemap coverage.
- Why
- Bing explicitly recommends the combined signals for AI-powered search.
- Reference
- Bing sitemaps in AI-powered search
Are target URLs checked with both indexed and Live URL views in Bing URL Inspection?
Priority: 5
- What
- Confirmed platform requirement: Separate stored index state from current fetch state.
- How
- inspect discovery, crawl, HTTP, HTML, canonical, markup, and live response.
- Why
- The tool shows what Bingbot sees and why a URL is excluded.
- Reference
- Bing URL Inspection
Is Bing crawl throttling configured through supported controls rather than error responses?
Priority: 4
- What
- Confirmed platform requirement: Protect capacity without destroying access.
- How
- use Crawl Control or a deliberate `crawl-delay`, then monitor freshness.
- Why
- Bing honors crawl-delay ahead of dashboard settings.
- Reference
- Bing Crawl Control
PERPLEXITY: VERIFIED 23 JULY 2026
Is PerplexityBot allowed where Perplexity search visibility is desired?
Priority: 5
- What
- Confirmed platform requirement: Permit the automatic search-index crawler.
- How
- test its robots group across canonical content and resources.
- Why
- Perplexity recommends allowing it to appear in results.
- Reference
- Perplexity crawler documentation
Does the WAF allow verified PerplexityBot traffic using both user-agent and IP?
Priority: 5
- What
- Confirmed platform requirement: Prevent security controls from nullifying robots intent.
- How
- combine UA matching with the official `perplexitybot.json` ranges.
- Why
- Perplexity explicitly recommends the combined condition.
- Reference
- Perplexity crawler documentation
Are Perplexity IP sets refreshed automatically from official endpoints?
Priority: 4
- What
- Confirmed platform requirement: Keep Cloudflare/AWS rules current.
- How
- periodically fetch both bot and user JSON sources, validate, stage, and alert on deltas.
- Why
- Perplexity says the addresses update regularly.
- Reference
- Perplexity crawler documentation
Is Perplexity-User treated separately from PerplexityBot?
Priority: 4
- What
- Confirmed platform requirement: Audit user-requested page fetches independently.
- How
- test both identities and document desired access.
- Why
- Perplexity-User supports question-time actions and is not the automatic index crawler.
- Reference
- Perplexity crawler documentation
Was Perplexity's stated policy-propagation delay allowed after crawler-rule changes?
Priority: 4
- What
- Confirmed platform requirement: Avoid testing too early.
- How
- repeat validation after the documented window and confirm logs.
- Why
- Perplexity says changes may take up to 24 hours.
- Reference
- Perplexity crawler documentation
Is PerplexityBot absent from the model-training opt-out matrix?
Priority: 5
- What
- Confirmed platform requirement: Do not mislabel its purpose.
- How
- classify it as search indexing and handle training policy elsewhere.
- Why
- Perplexity states the bot is not used to crawl for foundation models.
- Reference
- Perplexity crawler documentation
If all Perplexity mention is prohibited, is robots disallow supplemented by an appropriate removal control?
Priority: 4
- What
- Confirmed platform requirement: Account for residual domain/headline/summary indexing.
- How
- verify current displayed residue and escalate through publisher support where needed.
- Why
- Perplexity says blocked sites may still have limited facts indexed.
- Reference
- How Perplexity follows robots.txt
CLAUDE & ANTHROPIC: VERIFIED 23 JULY 2026
Is Claude-SearchBot allowed where Claude search visibility is desired?
Priority: 5
- What
- Confirmed platform requirement: Permit the search-quality crawler.
- How
- test the exact token across every host and key path.
- Why
- Anthropic says disabling it may reduce visibility and accuracy in user search results.
- Reference
- Anthropic crawler controls
Is Claude-User access deliberately configured and tested?
Priority: 4
- What
- Confirmed platform requirement: Decide whether Claude may retrieve pages at a user's direction.
- How
- test representative public, paywalled, and protected paths.
- Why
- Disabling it prevents user-initiated retrieval and may reduce visibility.
- Reference
- Anthropic crawler controls
Is ClaudeBot training policy independent from Claude-SearchBot and Claude-User?
Priority: 5
- What
- Confirmed platform requirement: Separate model-development collection from answer retrieval.
- How
- maintain three explicit robots groups.
- Why
- Anthropic documents a different purpose and effect for each.
- Reference
- Anthropic crawler controls
Are Anthropic rules present on every opted-in or opted-out subdomain?
Priority: 5
- What
- Confirmed platform requirement: Avoid partial policy coverage.
- How
- fetch `/robots.txt` and target responses host by host.
- Why
- Anthropic explicitly asks publishers to configure every subdomain.
- Reference
- Anthropic crawler controls
If Anthropic crawl load needs control, is its supported non-standard `crawl-delay` used carefully?
Priority: 3
- What
- Confirmed platform requirement: Throttle rather than block or error.
- How
- set a modest token-specific value and inspect freshness and logs.
- Why
- Anthropic says its bots support crawl-delay where appropriate.
- Reference
- Anthropic crawler controls
Are Anthropic bot source IPs verified against the current official range list?
Priority: 5
- What
- Confirmed platform requirement: Use Anthropic's published bot ranges as an authentication signal rather than assuming a static list.
- How
- fetch the official JSON on a controlled schedule, match the declared agent and source range, and alert on changes.
- Why
- Anthropic now publishes a living bot range list and warns that IP blocking is not a persistent opt-out method.
- Reference
- Anthropic bot IP ranges
Is `noindex` used when content must not appear in Claude web-search outputs?
Priority: 5
- What
- Confirmed platform requirement: Signal search partners not to index the page.
- How
- serve a crawlable `noindex` and verify disappearance over time.
- Why
- Anthropic documents this as an all-content-type exclusion control.
- Reference
- Anthropic blocking and removal
Is confidential content protected by authentication rather than crawler etiquette?
Priority: 5
- What
- Confirmed platform requirement: Enforce access at the application boundary.
- How
- require authorization and remove leaked public URLs. Use the owner-verified removal route if already surfaced.
- Why
- Anthropic recommends password protection for private material.
- Reference
- Anthropic blocking and removal
META AI: PRIMARY SOURCE VERIFIED 23 JULY 2026
Is Meta-WebIndexer explicitly allowed to crawl every public URL class intended for Meta AI discovery?
Priority: 5
- What
- Evidence-backed tactic: Maintain a token-specific robots policy for public canonical content.
- How
- Publish a `User-agent: meta-webindexer` group, allow intended public routes, exclude nonpublic routes, and compare the production robots response with request logs.
- Why
- Meta says this crawler supports Meta AI result relevance and accuracy and that allowing it helps Meta cite and link to content.
- Reference
- Meta Web Crawlers — Meta-WebIndexer
Is Meta-ExternalAgent governed separately from Meta-WebIndexer under the training and indexing policy?
Priority: 5
- What
- Evidence-backed tactic: Make an explicit allow-or-disallow decision for Meta-ExternalAgent.
- How
- Publish a separate `User-agent: meta-externalagent` group, document its owner and permitted URL scope, and review it independently from WebIndexer.
- Why
- Meta assigns ExternalAgent a different purpose: foundation-model training or product improvement through direct indexing.
Do private and sensitive routes remain protected if Meta-ExternalFetcher bypasses robots.txt?
Priority: 5
- What
- Evidence-backed tactic: Treat robots.txt as a crawl preference for user-requested fetches. It is not an authorization boundary.
- How
- Require authentication and server-side authorization for every nonpublic resource, remove secrets from public URLs and markup, and test protected routes while logged out.
- Why
- Meta says ExternalFetcher retrieves individual links at a user's request and may bypass robots.txt.
Are Meta robots.txt changes evaluated after the documented cache window?
Priority: 4
- What
- Evidence-backed tactic: Account for a propagation period of up to 24 hours.
- How
- Record publication time, avoid contradictory edits during the window, compare crawler requests before and after it, and do not declare a rule ineffective prematurely.
- Why
- Meta says its crawlers may cache robots.txt for up to 24 hours.
- Reference
- Meta Web Crawlers — robots.txt
Can logs distinguish Meta-WebIndexer, Meta-ExternalAgent, and Meta-ExternalFetcher traffic?
Priority: 4
- What
- Evidence-backed tactic: Preserve a normalized crawler identity beside the raw User-Agent.
- How
- Classify the documented tokens separately and retain time, URL, status, bytes, latency, source IP, and verification state.
- Why
- Meta documents different purposes and User-Agent strings for these crawlers.
Are requests claiming a Meta crawler identity verified before receiving WAF, rate-limit, or access exemptions?
Priority: 5
- What
- Evidence-backed tactic: Use User-Agent matching for classification without treating it as sufficient proof of trusted identity.
- How
- Match the documented token and use a Meta-published source-IP method only where the official page provides one. Do not borrow or invent ranges for newer agents.
- Why
- A request should not gain privileged treatment merely by presenting a documented User-Agent.
Can Meta crawler requests fetch public URLs without triggering writes, purchases, sessions, or other state changes?
Priority: 5
- What
- Evidence-backed tactic: Keep anonymously crawlable URLs read-only and safe to repeat.
- How
- Require authenticated intent, authorization, and confirmation for state-changing operations. Protect them against CSRF. Reject unsupported methods. Verify that anonymous GET and HEAD requests have no side effects.
- Why
- Meta says ExternalFetcher can support agentic site navigation for user tasks and may bypass robots.txt.
Are Meta crawler preferences expressed in robots.txt rather than relying on NoAI tags?
Priority: 5
- What
- Confirmed platform requirement: Encode Meta crawl choices through the relevant robots.txt agent groups.
- How
- Audit the production root robots file, add an explicit group for each applicable Meta token, validate the served response, and keep licensing notices separate from crawler enforcement.
- Why
- Meta identifies robots.txt as its supported industry-standard preference mechanism and contrasts it with nonstandard NoAI tags.
- Reference
- Meta Web Crawlers — robots.txt
APPLE INTELLIGENCE: VERIFIED 23 JULY 2026
Is Applebot allowed where Apple search and source-linked answer visibility is desired?
Priority: 5
- What
- Confirmed platform requirement: Permit Apple's search crawler.
- How
- audit Applebot's robots group, HTTP access, and logs.
- Why
- Apple says crawled data powers Spotlight, Siri, Safari, and current-context AI answers.
- Reference
- About Applebot
Is Applebot-Extended policy independent from Applebot search access?
Priority: 5
- What
- Confirmed platform requirement: Choose foundation-model training use separately.
- How
- configure the Extended token without blocking Applebot if search visibility remains desired.
- Why
- Applebot-Extended is a use-control token and does not crawl pages itself.
- Reference
- About Applebot
Is `nosnippet` used only when content should be excluded from Apple AI answer context?
Priority: 5
- What
- Confirmed platform requirement: Control answer grounding while preserving ordinary discoverability.
- How
- apply Applebot-specific meta or header rules to exact content.
- Why
- Apple says `nosnippet` opts content out of broad-world-knowledge AI answers.
- Reference
- About Applebot
Is paywalled or metered content marked `isAccessibleForFree: false` for Applebot?
Priority: 5
- What
- Confirmed platform requirement: Declare page-level restricted access.
- How
- add accurate JSON-LD and validate it against the served page.
- Why
- Apple keeps marked pages eligible for results but excludes their content from AI answer context.
- Reference
- About Applebot
Are Applebot controls for PDFs, images, and other non-HTML resources sent through `X-Robots-Tag`?
Priority: 4
- What
- Confirmed platform requirement: Cover assets that cannot contain HTML meta tags.
- How
- inspect headers at CDN and origin for each file type.
- Why
- Apple documents header-level directives for non-HTML resources.
- Reference
- About Applebot
Does robots.txt explicitly mention Applebot rather than accidentally inheriting Googlebot rules?
Priority: 4
- What
- Confirmed platform requirement: Remove ambiguous fallback behavior.
- How
- add a deliberate Applebot group and test it.
- Why
- Apple says it follows Googlebot instructions when Applebot is absent but Googlebot is present.
- Reference
- About Applebot
Are Applebot search, answer-context, and model-training outcomes tested as three separate states?
Priority: 4
- What
- Evidence-backed tactic: Verify the intended combination.
- How
- check search discoverability, snippet/context directives, and Applebot-Extended policy independently.
- Why
- Apple's documented controls affect different product uses.
- Reference
- About Applebot
Is Applebot traffic authenticated using Apple's documented verification approach before WAF decisions?
Priority: 4
- What
- Confirmed platform requirement: Distinguish real traffic from spoofed UA strings.
- How
- follow Apple's bot-verification instructions and retain evidence in security logs.
- Why
- access decisions based on a bare UA are unsafe.
- Reference
- About Applebot
NETWORK ACCESS & BOT DELIVERY
Do eligible pages return a real `200` response with substantive content?
Priority: 5
- What
- Confirmed platform requirement: Verify status and body together.
- How
- fetch from origin and edge as each target bot.
- Why
- A success code sends content to processing, but empty/error bodies can become soft 404s.
- Reference
- Google HTTP status handling
Are application error pages prevented from returning `200`?
Priority: 5
- What
- Confirmed platform requirement: Eliminate soft 404s.
- How
- map missing, expired, and failed records to meaningful statuses and test SPA fallbacks.
- Why
- Crawlers may treat error-like `200` content as absent.
- Reference
- Google JavaScript SEO basics
Do removed resources return `404` or `410` rather than a thin success page?
Priority: 5
- What
- Confirmed platform requirement: Communicate permanent absence.
- How
- test deleted URLs through CDN, redirects, and application routes.
- Why
- Correct status accelerates removal and avoids stale retrieval.
- Reference
- Google HTTP status handling
Do permanent moves use a direct permanent redirect with a short chain?
Priority: 5
- What
- Confirmed platform requirement: Consolidate old URLs to the final canonical.
- How
- test every hop, status, and target response.
- Why
- Permanent redirects are canonical signals and long chains waste crawl work.
- Reference
- Google redirects guidance
Are `429` and `5xx` responses rare, monitored, and correctly scoped?
Priority: 5
- What
- Confirmed platform requirement: Prevent crawl throttling and eventual index loss.
- How
- alert by bot, route, edge, and origin. Repair capacity rather than using errors as routine rate control.
- Why
- Google slows crawling on `429`/`5xx` and may eventually drop persistent failures.
- Reference
- Google HTTP status handling
Is `/robots.txt` highly available with intentional status behavior?
Priority: 5
- What
- Confirmed platform requirement: Treat it as critical infrastructure.
- How
- bypass brittle application dependencies, monitor content/status, and test failure modes.
- Why
- A robots `5xx` can stop Google crawling while it retries.
- Reference
- Google robots.txt specification
Are DNS, TLS, IPv4/IPv6, and edge routing healthy from external bot regions?
Priority: 5
- What
- Evidence-backed tactic: Verify the connection before HTML concerns.
- How
- synthetic-probe resolution, certificate chain, handshake, and response from multiple geographies.
- Why
- Network failures prevent any retrieval.
- Reference
- Google network and DNS errors
Are intended bots exempt from CAPTCHA, JavaScript challenge, interstitial, and login loops?
Priority: 5
- What
- Confirmed platform requirement: Remove non-content gates from crawler paths.
- How
- combine verified identity with narrowly scoped WAF rules and test clean sessions.
- Why
- OpenAI identifies these controls as common crawler failures.
- Reference
- OpenAI advertiser crawler guidance
Are country and ASN blocks checked against crawler egress locations?
Priority: 5
- What
- Evidence-backed tactic: Prevent invisible geo-denial.
- How
- test US and other documented crawler locations through the complete edge policy.
- Why
- Google primarily egresses from the US and may crawl from elsewhere.
- Reference
- Google crawler overview
Are `ETag`, `Last-Modified`, and correct `304` responses implemented for stable resources?
Priority: 3
- What
- Evidence-backed tactic: Enable efficient recrawling.
- How
- validate conditional requests and ensure updates invalidate validators.
- Why
- Google supports both cache validators and can reuse unchanged content.
- Reference
- Google crawler overview
Does critical text occur before crawler file-size cutoffs?
Priority: 4
- What
- Confirmed platform requirement: Keep primary content and metadata early and responses lean.
- How
- measure uncompressed HTML/resource sizes and test truncation.
- Why
- Google crawlers stop processing beyond product-specific limits.
- Reference
- Google crawler overview
Do APIs, media, CSS, and JavaScript return correct MIME types and unrestricted bot responses?
Priority: 4
- What
- Evidence-backed tactic: Ensure dependencies are fetchable and interpretable.
- How
- crawl resource graphs and compare status, content type, cache, and WAF results.
- Why
- Blocked or malformed resources can hide meaning and markup.
- Reference
- Google JavaScript SEO basics
Are search, training, and user-triggered fetchers monitored as separate agents?
Priority: 5
- What
- Evidence-backed tactic: Preserve user agent, verified source, URL, status, bytes, and timing.
- How
- Create purpose-specific log views rather than one “AI bot” bucket.
- Why
- Access decisions and diagnostic meaning differ by crawler role.
- Reference
- Anthropic crawler documentation
Are claimed AI crawlers validated with official IP sources or documented verification rather than user agent alone?
Priority: 5
- What
- Evidence-backed tactic: Separate genuine platform access from spoofed bot traffic.
- How
- Match user agent and current official IP data, retaining verification status.
- Why
- User-agent strings are easy to spoof.
- Reference
- Perplexity crawler documentation
Are ChatGPT search inclusion and model-training preferences configured and audited independently?
Priority: 5
- What
- Confirmed platform requirement: Treat OAI-SearchBot and GPTBot as different controls.
- How
- Review robots rules and logs for each agent after changes.
- Why
- OpenAI states that search inclusion and training opt-out use separate bots.
- Reference
- OpenAI publishers and developers FAQ
Are `PerplexityBot` indexing requests distinguished from `Perplexity-User` user-triggered fetches?
Priority: 5
- What
- Confirmed platform requirement: Monitor and configure the two roles independently.
- How
- Use official user agents and current published IP ranges in WAF and logs.
- Why
- Perplexity documents different purposes and robots behavior for the two agents.
- Reference
- Perplexity crawler documentation
Are `ClaudeBot`, `Claude-User`, and other documented Anthropic agents configured according to intended use?
Priority: 4
- What
- Confirmed platform requirement: Avoid using one robots rule as a proxy for all Anthropic access.
- How
- Review the living crawler documentation before each policy change.
- Why
- Training collection and user-requested access are distinct purposes.
- Reference
- Anthropic crawler documentation
Are verified answer-engine requests monitored for `403`, challenge, rate-limit, and origin errors?
Priority: 5
- What
- Evidence-backed tactic: Detect accidental denial outside robots.txt.
- How
- Alert on status shifts by verified agent and sample affected URLs.
- Why
- A permissive robots file does not override a blocking WAF or CDN.
- Reference
- Perplexity crawler documentation
Are representative public URLs automatically checked for robots, indexability, renderability, and fetch status after deployments?
Priority: 4
- What
- Evidence-backed tactic: Detect access regressions before dashboard trends decline.
- How
- Test an English/Turkish canary set through public paths and verified log evidence.
- Why
- Small configuration changes can block an entire content class.
- Reference
- Google AI features and your website
When using `noindex`, can the intended crawler fetch the page and read the directive?
Priority: 5
- What
- Confirmed platform requirement: Avoid blocking the crawler before it sees the exclusion signal.
- How
- Test robots access and rendered meta or header directives together.
- Why
- OpenAI notes that its crawler must access the page to detect `noindex`.
- Reference
- OpenAI publishers and developers FAQ
If Cloudflare Markdown for Agents is enabled, does the origin send an explicit `Content-Signal` header that matches the approved policy?
Priority: 4
- What
- Confirmed platform requirement: Prevent converted Markdown responses from silently receiving Cloudflare's permissive default.
- How
- Set the header at the origin, then request representative URLs with `Accept: text/markdown` and verify the effective response across routes and caches.
- Why
- Cloudflare preserves an origin header, but otherwise adds `ai-train=yes, search=yes, ai-input=yes` to converted Markdown.
RENDERING, LINKS & CANONICALS
Is the primary answer-bearing text present in the initial HTML?
Priority: 5
- What
- Evidence-backed tactic: Avoid making retrieval depend entirely on client execution.
- How
- compare raw response with rendered DOM for headings, facts, links, and citations.
- Why
- Google recommends server/pre-rendering and notes that not all bots run JavaScript.
- Reference
- Google JavaScript SEO basics
Does Google's rendered HTML contain the same intended content and directives as the source response?
Priority: 5
- What
- Confirmed platform requirement: Detect hydration and rendering loss.
- How
- use URL Inspection rendered HTML and screenshot, then compare canonical, robots, structured data, and body.
- Why
- Google indexes rendered HTML for JavaScript pages.
- Reference
- Google JavaScript SEO basics
Is primary content available without clicks, swipes, typing, or consent interaction?
Priority: 5
- What
- Confirmed platform requirement: Remove interaction-gated information from the retrieval path.
- How
- inspect clean-session HTML and rendered DOM before any action.
- Why
- Google does not trigger user interaction to lazy-load primary content.
- Reference
- Mobile-first indexing best practices
Are render-critical JavaScript and CSS paths crawlable?
Priority: 5
- What
- Confirmed platform requirement: Prevent incomplete rendering.
- How
- crawl the resource graph and inspect robots rules, responses, and CSP/CDN failures.
- Why
- Google will not render JavaScript from blocked files or pages.
- Reference
- Google JavaScript SEO basics
Are discoverable links rendered as real crawlable anchors with resolvable URLs?
Priority: 5
- What
- Confirmed platform requirement: Expose navigation and citations in standard HTML.
- How
- inspect raw/rendered anchors and crawl from hubs without a browser history API.
- Why
- Crawlers primarily discover URLs through links.
- Reference
- Googlebot documentation
Does every indexable page declare one valid canonical matching the intended public URL?
Priority: 5
- What
- Confirmed platform requirement: Consolidate duplicates.
- How
- validate status, absolute URL, indexability, and reciprocal internal signals.
- Why
- Canonicalization decides which version can carry index and citation signals.
- Reference
- Google canonical guidance
Do redirects, HTML/header canonicals, sitemap URLs, hreflang, and internal links agree?
Priority: 5
- What
- Evidence-backed tactic: Remove contradictory URL identity signals.
- How
- build a per-page signal matrix and fail conflicting targets.
- Why
- Consistency improves consolidation and crawl efficiency.
- Reference
- Google canonical guidance
Are tracking, sort, filter, print, session, and case variants prevented from becoming competing answer candidates?
Priority: 4
- What
- Evidence-backed tactic: Control duplicate URL inventory.
- How
- normalize links, redirect true duplicates, canonicalize when needed, and block useless crawl spaces.
- Why
- Duplicate inventory consumes crawl resources and fragments identity.
- Reference
- Google crawl budget guidance
Do canonical signals cover PDFs and other non-HTML documents?
Priority: 4
- What
- Confirmed platform requirement: Consolidate equivalent downloadable and HTML versions.
- How
- serve a `Link: <...>; rel="canonical"` response header where appropriate.
- Why
- HTML canonical tags cannot live inside non-HTML files.
- Reference
- Google canonical guidance
Are mobile and desktop robots directives, title, description, canonicals, and structured data equivalent?
Priority: 5
- What
- Confirmed platform requirement: Avoid device-specific eligibility loss.
- How
- compare served and rendered variants.
- Why
- Google indexes the mobile version and warns that divergent directives can block indexing.
- Reference
- Mobile-first indexing best practices
Does each citation-worthy document and media asset have a stable, directly fetchable URL?
Priority: 4
- What
- Evidence-backed tactic: Avoid transient blobs, expiring signed links, or app-only destinations.
- How
- test long-lived URLs without session state.
- Why
- Answer systems need resolvable sources to retrieve and link.
- Reference
- OpenAI advertiser crawler guidance
Is page structure encoded with semantic headings, landmarks, accessible names, roles, and states?
Priority: 3
- What
- Evidence-backed tactic: Make content and actions machine-readable.
- How
- run accessibility-tree and agent-task audits.
- Why
- OpenAI says Atlas uses ARIA to interpret interactive pages.
- Reference
- OpenAI publisher FAQ
Are indexable routes free of hash-fragment-only content addressing?
Priority: 4
- What
- Confirmed platform requirement: Give content stable server-resolvable paths.
- How
- request every canonical URL without prior app state and migrate legacy AJAX fragments.
- Why
- Google deprecated the AJAX crawling scheme and does not rely on URL fragments for separate documents.
- Reference
- Fix JavaScript search problems
Is structured data present and valid in the rendered page after deployment?
Priority: 5
- What
- Confirmed platform requirement: Test the output crawlers receive alongside the templates.
- How
- run Rich Results Test, URL Inspection, and Bing markup inspection on production URLs.
- Why
- templating or serving failures can break valid source code after release.
- Reference
- Google structured data introduction
SITEMAPS, FEEDS & FRESHNESS
Does the XML sitemap contain only preferred canonical, indexable production URLs?
Priority: 5
- What
- Confirmed platform requirement: Publish a clean discovery inventory.
- How
- diff sitemap entries against canonicals, status, noindex, environment, and redirects.
- Why
- Search engines use sitemaps as canonical and discovery hints.
- Reference
- Google sitemap guidance
Are sitemap locations fully qualified, absolute, encoded URLs that return the intended resource?
Priority: 5
- What
- Confirmed platform requirement: Eliminate malformed discovery targets.
- How
- parse, normalize, fetch, and compare each sampled URL.
- Why
- Google crawls URLs exactly as listed.
- Reference
- Google sitemap guidance
Are large sitemaps split within protocol limits and referenced by valid indexes?
Priority: 4
- What
- Confirmed platform requirement: Keep every file processable.
- How
- enforce at most 50 MB uncompressed or 50,000 URLs per sitemap and validate indexes.
- Why
- Oversized files can fail processing.
- Reference
- Google sitemap guidance
Does `lastmod` change only after a significant page update?
Priority: 5
- What
- Confirmed platform requirement: Send truthful freshness signals.
- How
- derive it from main content, structured data, or meaningful link changes. Do not use deploy or sitemap time.
- Why
- Bing uses accurate values to focus crawling.
- Reference
- Bing sitemaps in AI-powered search
Are sitemap `priority` and `changefreq` excluded from GEO scoring and operational promises?
Priority: 3
- What
- Confirmed platform requirement: Stop relying on ignored fields.
- How
- remove dashboards and checks that treat them as freshness/rank inputs.
- Why
- Bing says both fields are ignored.
- Reference
- Bing sitemaps in AI-powered search
Are sitemap submission status, last-read time, and processing errors monitored?
Priority: 5
- What
- Confirmed platform requirement: Confirm that discovery infrastructure is actually consumed.
- How
- alert on stale reads, fetch failures, and parse errors in Google and Bing webmaster tools.
- Why
- A published but unread sitemap provides no timely signal.
- Reference
- Bing sitemaps in AI-powered search
Where IndexNow is used, does the pipeline handle publish, substantive update, redirect, and deletion events without duplicate storms?
Priority: 2
- What
- Evidence-backed tactic: Operate optional change notifications as a reliable, idempotent freshness channel.
- How
- Queue canonical URLs, debounce cosmetic changes, retry from response codes, log receipts, and reconcile missed events.
- Why
- IndexNow recommends automated lifecycle notifications while discouraging repeat submissions without meaningful changes.
- Reference
- IndexNow operational FAQ
Are RSS/Atom feeds available and submitted for rapidly changing editorial content?
Priority: 4
- What
- Evidence-backed tactic: Provide a recent-change stream alongside the full sitemap.
- How
- validate entries, canonical links, timestamps, and feed fetchability.
- Why
- Google accepts RSS 2.0 and Atom 1.0 as sitemap formats.
- Reference
- Google sitemap guidance
Has the retired Google sitemap ping endpoint been removed?
Priority: 3
- What
- Confirmed platform requirement: Stop sending no-op HTTP pings.
- How
- submit through robots.txt, Search Console/API, sitemap fetch, RSS/Atom, or WebSub as appropriate.
- Why
- Google deprecated the ping endpoint and it returns `404`.
- Reference
- Google sitemap ping deprecation
Do direct product feeds and public landing pages update atomically?
Priority: 4
- What
- Evidence-backed tactic: Prevent an answer engine from seeing mismatched availability, price, URL, or description.
- How
- publish with versioned jobs, validate both surfaces, then notify discovery systems.
- Why
- OpenAI notes delays and recommends direct feeds for current product data.
- Reference
- Shopping with ChatGPT Search
MEASUREMENT DESIGN
Is every GEO measurement tied to a specific decision rather than a generic visibility score?
Priority: 5
- What
- Evidence-backed tactic: State the decision, audience, metric, decision threshold, and owner.
- How
- Put these fields at the top of every recurring report and experiment brief.
- Why
- Metrics without a decision invite post-hoc narratives and metric shopping.
- Reference
- OpenAI evaluation best practices
Does each GEO metric have a named owner, review cadence, and escalation recipient?
Priority: 4
- What
- Evidence-backed tactic: Record accountability for collection, interpretation, and remediation.
- How
- Maintain a lightweight RACI beside the metric dictionary.
- Why
- Unowned dashboards decay and alerts go unanswered.
- Reference
- NIST AI 600-1 Generative AI Profile
Are answer mentions, citations, referral visits, conversions, and revenue reported as separate outcome layers?
Priority: 5
- What
- Evidence-backed tactic: Use a funnel with distinct denominators instead of one blended GEO score.
- How
- Define a metric for each observable layer and disclose missing links between them.
- Why
- A citation is neither a click nor a conversion.
Are GEO metrics defined with numerator, denominator, unit, scope, and known blind spots?
Priority: 4
- What
- Evidence-backed tactic: Document exact formulas for prevalence, citation share, accuracy, and conversion metrics.
- How
- Version a shared metric dictionary and link it from every dashboard.
- Why
- Similar labels often hide incompatible calculations.
Is the exact set of answer engines, surfaces, accounts, and markets in scope recorded?
Priority: 5
- What
- Evidence-backed tactic: Name each product surface rather than grouping everything under “AI search.”
- How
- Track platform, surface, access mode, subscription state, language, and country.
- Why
- Products can retrieve, cite, personalize, and report data differently.
- Reference
- ChatGPT Search help
Does monitoring include the brand's official name, spelling variants, products, people, and common aliases?
Priority: 5
- What
- Evidence-backed tactic: Maintain the entity strings that count as a mention and those that create ambiguity.
- How
- Review aliases with brand, legal, and local-language owners quarterly.
- Why
- Exact-string matching undercounts variants and can merge unrelated entities.
- Reference
- NIST AI 600-1 Generative AI Profile
Is the owned, operated, partner, reseller, and unaffiliated domain set explicitly classified?
Priority: 4
- What
- Evidence-backed tactic: Establish which cited hosts count as first-party visibility.
- How
- Keep a versioned registry including subdomains, migrations, and country domains.
- Why
- Host-level aggregation can misattribute partner or legacy properties.
Is the competitor or peer set versioned with clear inclusion rules?
Priority: 3
- What
- Evidence-backed tactic: Preserve the entities used in share and recommendation comparisons.
- How
- Add or remove peers only at declared reporting boundaries.
- Why
- Changing the peer set changes normalized shares even if engine behavior is unchanged.
Are start time, end time, timezone, cadence, and blackout periods recorded?
Priority: 5
- What
- Evidence-backed tactic: Make the temporal sampling frame explicit.
- How
- Store UTC timestamps and a business timezone for every run.
- Why
- Engines and indexed content change over time.
Is each observation accompanied by platform, surface, prompt, context, locale, account state, and timestamp?
Priority: 5
- What
- Evidence-backed tactic: Store enough metadata to explain or reproduce a run.
- How
- Use an immutable run ID and structured observation schema.
- Why
- A response without execution context is weak evidence.
Are full response text, citations, visible source panels, and screenshots retained subject to policy?
Priority: 5
- What
- Evidence-backed tactic: Keep auditable evidence behind derived scores.
- How
- Store a timestamped snapshot with access controls and retention limits.
- Why
- Interfaces and citations can change after collection.
- Reference
- NIST AI 600-1 Generative AI Profile
PROMPT CORPUS & OBSERVATION CAPTURE
Are direct brand and navigational prompts measured separately from discovery prompts?
Priority: 4
- What
- Evidence-backed tactic: Create a dedicated segment for users already seeking the entity.
- How
- Tag brand-only, brand-plus-topic, and URL/navigation prompts.
- Why
- High visibility on navigational queries can mask weak category discovery.
- Reference
- OpenAI evaluation best practices
Does the prompt set include unbranded category and problem-discovery needs?
Priority: 5
- What
- Evidence-backed tactic: Measure whether the brand appears before it is named.
- How
- Source prompts from search demand, support logs, sales calls, and user research.
- Why
- Unbranded discovery is a different exposure opportunity from brand recall.
- Reference
- OpenAI evaluation best practices
Are “best,” “alternative,” “versus,” and shortlist prompts analyzed as a distinct segment?
Priority: 5
- What
- Evidence-backed tactic: Track inclusion, exclusion, stated criteria, and cited evidence in comparative answers.
- How
- Build balanced prompts across competitors and decision criteria.
- Why
- Comparative answers carry higher commercial and brand-risk stakes.
Are informational prompts that should cite first-party facts measured separately from recommendation prompts?
Priority: 4
- What
- Evidence-backed tactic: Identify questions about specifications, policies, people, prices, and processes.
- How
- Map each prompt to a maintained source-of-truth URL and expected fact set.
- Why
- Factual accuracy and commercial recommendation require different rubrics.
Is each prompt tagged by awareness, consideration, decision, support, or retention intent?
Priority: 3
- What
- Evidence-backed tactic: Connect visibility to the expected next user action.
- How
- Apply a documented intent rubric and double-review ambiguous prompts.
- Why
- Equal weighting across unlike intents distorts business significance.
- Reference
- OpenAI evaluation best practices
Are English and Turkish prompt results collected and reported as separate cohorts?
Priority: 5
- What
- Evidence-backed tactic: Preserve language-specific exposure, accuracy, and citation metrics.
- How
- Use native prompts rather than literal translations and compare only equivalent intents.
- Why
- Aggregation can hide language-specific failure and source selection.
- Reference
- NIST AI 600-1 Generative AI Profile
Are country, interface locale, and query language recorded independently?
Priority: 4
- What
- Evidence-backed tactic: Distinguish geography from language in every observation.
- How
- Record declared location, observed location, locale, timezone, and VPN/proxy use.
- Why
- Search results and query rewrites can vary with general location.
- Reference
- ChatGPT Search help
Is logged-in, memory-enabled, or personalized testing separated from clean-session testing?
Priority: 4
- What
- Evidence-backed tactic: Treat user context as an experimental factor.
- How
- Record account state and run a stable clean-session benchmark alongside real-user scenarios.
- Why
- Memory and conversation context can alter query rewriting and responses.
- Reference
- ChatGPT Search help
Are desktop, mobile, app, browser, and embedded surfaces identified in observations?
Priority: 3
- What
- Evidence-backed tactic: Preserve the interface through which each answer was generated.
- How
- Record device class and product surface with each run.
- Why
- Available presentation and local features can vary by surface.
- Reference
- ChatGPT Search help
Is the benchmark grounded in real user needs rather than only synthetic prompts?
Priority: 5
- What
- Evidence-backed tactic: Blend consented production questions, research, search demand, and expert-authored edge cases.
- How
- Record provenance and sampling rules for every prompt.
- Why
- A convenient prompt list can measure the benchmark designer rather than the audience.
- Reference
- OpenAI evaluation best practices
Does the reporting retain real prompt-frequency weights while also showing an unweighted diagnostic view?
Priority: 4
- What
- Evidence-backed tactic: Separate business-weighted performance from coverage across unique needs.
- How
- Store both the occurrence weight and one-vote-per-prompt result.
- Why
- Deduplicating all repeated needs can erase demand concentration.
Does the corpus include ambiguous, misspelled, adversarial, and high-risk fact prompts?
Priority: 4
- What
- Evidence-backed tactic: Test realistic failure modes beyond the happy path.
- How
- Maintain an edge-case suite informed by incidents, support tickets, and red-team review.
- Why
- Average visibility can coexist with severe brand-safety failures.
- Reference
- OpenAI evaluation best practices
Are personal, confidential, contractual, and embargoed details removed before prompts reach external systems?
Priority: 5
- What
- Evidence-backed tactic: Apply data minimization and approval rules to the test corpus.
- How
- Redact identifiers, use synthetic substitutes, and prohibit secret-bearing prompts.
- Why
- GEO monitoring should not become an uncontrolled disclosure channel.
- Reference
- NIST AI 600-1 Generative AI Profile
Is every benchmark run tied to an immutable prompt-set version?
Priority: 5
- What
- Evidence-backed tactic: Preserve prompt text, metadata, weights, additions, removals, and rationale.
- How
- Use a versioned repository or append-only dataset manifest.
- Why
- Trend lines are uninterpretable when the denominator silently changes.
- Reference
- NIST AI 600-1 Generative AI Profile
Is a stable subset of prompts retained for longitudinal comparison?
Priority: 4
- What
- Evidence-backed tactic: Keep a frozen panel while allowing a separate evolving discovery corpus.
- How
- Version both and never splice their trend lines without restatement.
- Why
- Corpus refreshes otherwise masquerade as performance change.
Is an unseen prompt subset reserved to test whether an optimization generalizes?
Priority: 4
- What
- Evidence-backed tactic: Separate tuning prompts from validation prompts.
- How
- Restrict access to the holdout and evaluate it only at planned gates.
- Why
- Repeatedly optimizing on the same prompts overfits the benchmark.
- Reference
- OpenAI evaluation best practices
Are inclusion, exclusion, deduplication, and retirement rules documented?
Priority: 3
- What
- Evidence-backed tactic: Make corpus curation reproducible.
- How
- Require a reason code and reviewer for every material corpus change.
- Why
- Selective prompt removal can manufacture an improvement.
- Reference
- NIST AI 600-1 Generative AI Profile
Does each factual prompt identify the authoritative owned page or record expected to support it?
Priority: 5
- What
- Evidence-backed tactic: Create a prompt-to-evidence mapping.
- How
- Store expected URLs, approved facts, freshness owner, and review date.
- Why
- A brand cannot assess answer correctness without a maintained reference point.
- Reference
- NIST AI 600-1 Generative AI Profile
Are accuracy, relevance, tone, citation, and safety criteria specified before scoring?
Priority: 4
- What
- Evidence-backed tactic: Use a multi-dimensional rubric rather than a single “good answer” judgment.
- How
- Define anchored scales and unacceptable-failure conditions for each prompt class.
- Why
- Vague criteria produce inconsistent review and retrospective scoring.
Is the cited page content captured or hashed at observation time?
Priority: 4
- What
- Evidence-backed tactic: Preserve evidence of what the engine could have retrieved then.
- How
- Store a permitted snapshot or content checksum plus fetch timestamp and status.
- Why
- Later page edits can be mistaken for engine inconsistency.
Are timeouts, captchas, no-search responses, parsing failures, and partial answers retained as outcomes?
Priority: 4
- What
- Evidence-backed tactic: Make missingness visible instead of silently dropping failed runs.
- How
- Use explicit failure reason codes and report the failure rate.
- Why
- Non-random missing observations can bias visibility upward.
- Reference
- OpenAI evaluation best practices
Does measurement distinguish answers that searched the web from answers that did not?
Priority: 4
- What
- Evidence-backed tactic: Record observable search activation and citation availability.
- How
- Use interface indicators where available and label uncertain cases as unknown.
- Why
- Absence of a citation is not comparable when retrieval never occurred.
- Reference
- ChatGPT Search help
Are raw observations retained alongside deduplicated and normalized metrics?
Priority: 4
- What
- Evidence-backed tactic: Preserve citation events before host, page, or entity aggregation.
- How
- Build transformations from immutable raw records with versioned code.
- Why
- Aggregation rules can change conclusions and must be auditable.
Are browser, collector, extraction, classifier, and rubric versions attached to every run?
Priority: 4
- What
- Evidence-backed tactic: Treat measurement code changes as potential discontinuities.
- How
- Store commit or release identifiers and annotate deployments.
- Why
- A parser change can look like a visibility change.
- Reference
- NIST AI 600-1 Generative AI Profile
SAMPLING & STATISTICAL INTEGRITY
Is every important prompt sampled repeatedly rather than treated as a single rank check?
Priority: 5
- What
- Evidence-backed tactic: Estimate the response distribution for the same prompt and context.
- How
- Repeat at precommitted times and report the number of valid runs.
- Why
- Identical prompts can return different answers and citations.
Are observations spread across multiple days and relevant dayparts?
Priority: 5
- What
- Evidence-backed tactic: Capture temporal variability rather than a single collection burst.
- How
- Use a scheduled panel with timestamps and a stable prompt mix.
- Why
- Same-session bursts can understate non-stationarity.
Is prompt execution order randomized or counterbalanced?
Priority: 3
- What
- Evidence-backed tactic: Prevent time, throttling, or session-order effects from aligning with one treatment.
- How
- Randomize within blocks of platform, locale, and intent.
- Why
- Fixed order can confound treatment with collection conditions.
- Reference
- NIST AI 600-1 Generative AI Profile
Is sample size chosen before looking at favorable results?
Priority: 5
- What
- Evidence-backed tactic: Base collection volume on pilot variance, desired precision, and cost.
- How
- Write the stopping rule in the study plan and keep it fixed.
- Why
- Optional stopping inflates false discoveries.
Does collection continue to the precommitted endpoint even after a favorable result appears?
Priority: 5
- What
- Evidence-backed tactic: Enforce the planned stopping rule.
- How
- Lock the schedule or require documented independent approval to stop early.
- Why
- Stopping on a lucky result manufactures confidence.
Are visibility estimates reported with uncertainty intervals rather than point estimates alone?
Priority: 5
- What
- Evidence-backed tactic: Pair prevalence, share, and quality metrics with interval estimates.
- How
- Use an appropriate bootstrap or model and disclose assumptions.
- Why
- Small apparent differences may be measurement noise.
Does every score show prompts attempted, valid responses, and observations contributing to it?
Priority: 5
- What
- Evidence-backed tactic: Expose denominators and missing data.
- How
- Display `n` beside each segment and suppress underpowered comparisons.
- Why
- A percentage without its denominator looks more certain than it is.
Are low-frequency domains and brands flagged as especially unstable?
Priority: 4
- What
- Evidence-backed tactic: Treat sparse citation prevalence cautiously.
- How
- Show zero counts, intervals, and a low-base warning rather than forcing ranks.
- Why
- Rare events produce volatile percentages and ordinal positions.
Is platform drift tested before combining old and new observations?
Priority: 4
- What
- Evidence-backed tactic: Look for level, variance, or source-distribution changes over time.
- How
- Maintain rolling diagnostics and annotate platform or collector changes.
- Why
- Historical samples may no longer estimate the current system.
Does the methodology avoid declaring one fixed run count sufficient for every platform and prompt class?
Priority: 4
- What
- Evidence-backed tactic: Calibrate precision by local variance and decision risk.
- How
- Run a pilot, estimate uncertainty, then set a documented collection plan.
- Why
- Variability differs by platform, topic, language, and outcome.
Are treatment comparisons matched on platform, prompt, locale, surface, and collection window?
Priority: 5
- What
- Evidence-backed tactic: Block or pair observations on major context variables.
- How
- Use the same sampling schedule and prompt version for control and treatment.
- Why
- Cross-context differences can exceed the treatment effect.
- Reference
- NIST AI 600-1 Generative AI Profile
Are platform-specific metrics shown before any blended index?
Priority: 5
- What
- Evidence-backed tactic: Preserve each platform's distinct citation and retrieval behavior.
- How
- If a composite is used, disclose weights and show components.
- Why
- Raw citation counts are not directly comparable across platforms.
Is citation share normalized within a declared platform, prompt set, and peer universe?
Priority: 4
- What
- Evidence-backed tactic: State exactly what total forms the denominator.
- How
- Publish both raw citation count and normalized share.
- Why
- “Share of voice” is otherwise easy to misread as market share.
Do major conclusions hold under reasonable alternative prompt and business weights?
Priority: 3
- What
- Experimental practice: Test whether one weighting scheme drives the result.
- How
- Report weighted, unweighted, and segment-level views.
- Why
- Hidden weights can reverse comparative rankings.
Is the smallest decision-relevant improvement defined before an experiment?
Priority: 4
- What
- Evidence-backed tactic: Distinguish practical impact from any nonzero numerical change.
- How
- Set an effect threshold with stakeholders and power the study accordingly.
- Why
- Tiny changes can be statistically noisy or commercially irrelevant.
VISIBILITY & ANSWER QUALITY
Is brand mention prevalence measured as the share of valid responses containing the entity?
Priority: 5
- What
- Evidence-backed tactic: Distinguish response-level presence from citation count.
- How
- Apply the versioned alias set and publish the valid-response denominator.
- Why
- Repeated citations in one answer should not inflate how often the brand appears.
Is owned-domain citation prevalence tracked separately from brand mentions?
Priority: 5
- What
- Evidence-backed tactic: Count responses with at least one owned citation.
- How
- Resolve cited URLs against the versioned owned-domain registry.
- Why
- Engines can mention a brand without citing it, or cite it without naming it prominently.
Is raw citation count retained as a descriptive metric without being treated as a rank?
Priority: 3
- What
- Evidence-backed tactic: Count observed citation events before deduplication.
- How
- Report count beside prevalence, share, and sample size.
- Why
- Counts are useful operationally but depend heavily on response and platform format.
Is the number and distribution of unique owned pages cited tracked?
Priority: 4
- What
- Evidence-backed tactic: Measure whether visibility depends on one URL or a resilient content set.
- How
- Canonicalize URLs and report concentration by page.
- Why
- A domain total can hide fragile dependence on a single asset.
Is the diversity and concentration of all cited source domains monitored?
Priority: 3
- What
- Evidence-backed tactic: Track how concentrated answer evidence is across publishers.
- How
- Report unique domains and a concentration statistic by prompt segment.
- Why
- Source concentration can expose dependency and information-quality risks.
- Reference
- Synthetic Sources?
Are cited owned pages classified as product, service, research, guide, profile, policy, or support content?
Priority: 3
- What
- Experimental practice: Identify which asset types engines use for each intent.
- How
- Map canonical URLs to a stable content taxonomy.
- Why
- Page-type patterns can guide testing without pretending to reveal a ranking factor.
Is the entity's narrative prominence scored separately from citation occurrence?
Priority: 3
- What
- Experimental practice: Assess whether the brand is central, supporting, incidental, or absent.
- How
- Use an anchored human rubric and preserve the answer text.
- Why
- A footnote-level source and a central recommendation create different exposure.
- Reference
- GEO: Generative Engine Optimization
Does reporting avoid treating Bing citation totals or URL counts as placement data?
Priority: 5
- What
- Confirmed platform requirement: Label placement as unavailable unless directly observed and preserved.
- How
- Remove rank-like labels from Bing AI Performance exports.
- Why
- The platform explicitly says these metrics do not show placement or importance.
Is positive, neutral, negative, mixed, and purely factual treatment scored with an anchored rubric?
Priority: 4
- What
- Experimental practice: Separate visibility from the answer's stance toward the entity.
- How
- Use human-reviewed examples and an explicit “not applicable” state.
- Why
- More mentions can be harmful when the context is inaccurate or adverse.
Is shortlist or recommendation inclusion measured only for prompts where a recommendation is appropriate?
Priority: 5
- What
- Experimental practice: Record included, excluded, warned against, or not applicable.
- How
- Apply intent-specific scoring and retain the stated selection criteria.
- Why
- Treating every mention as a recommendation overstates commercial exposure.
- Reference
- GEO: Generative Engine Optimization
Is the cited page directly relevant to the user's question and cited claim?
Priority: 4
- What
- Evidence-backed tactic: Distinguish topical overlap from evidence that answers the claim.
- How
- Use an anchored relevance scale with examples.
- Why
- Broadly related sources can create a false appearance of support.
Are brand, product, people, policy, location, and pricing facts checked against current authoritative records?
Priority: 5
- What
- Evidence-backed tactic: Score fact-level correctness and materiality.
- How
- Maintain a dated truth set and require subject-matter review for high-risk claims.
- Why
- Citation presence does not guarantee the answer is correct.
- Reference
- NIST AI 600-1 Generative AI Profile
Are responses checked for merging the brand with similarly named organizations, people, or products?
Priority: 5
- What
- Evidence-backed tactic: Record identity collisions as a distinct critical error.
- How
- Test aliases and ambiguous prompts. Verify official identifiers and domains.
- Why
- An answer can appear visible while describing the wrong entity.
- Reference
- NIST AI 600-1 Generative AI Profile
Are claims credited to the correct organization, author, study, or page?
Priority: 5
- What
- Evidence-backed tactic: Score attribution independently from factual truth.
- How
- Compare the answer's attribution with the cited source's authorship and context.
- Why
- Correct facts can still damage trust when assigned to the wrong source.
Are cited facts and pages checked for current validity at collection time?
Priority: 5
- What
- Evidence-backed tactic: Track source date, last verified date, and whether the claim is time-sensitive.
- How
- Apply shorter review windows to pricing, people, policies, and regulatory facts.
- Why
- A well-cited answer can still be stale.
Is the share of citations to primary, official, or original evidence tracked for high-risk prompts?
Priority: 4
- What
- Evidence-backed tactic: Classify sources by provenance rather than domain popularity alone.
- How
- Use a documented source-tier rubric and review exceptions.
- Why
- Secondary summaries can introduce drift or strip conditions from claims.
- Reference
- NIST AI 600-1 Generative AI Profile
Are high-risk answers reviewed for reliance on apparently synthetic or low-accountability sources?
Priority: 3
- What
- Experimental practice: Examine source provenance, authorship, evidence chain, and editorial accountability.
- How
- Manually audit a risk-based sample and record uncertainty.
- Why
- Synthetic sources can amplify unsupported claims through recitation.
- Reference
- Synthetic Sources?
Is no source labeled AI-generated solely from an AI-content detector score?
Priority: 5
- What
- Evidence-backed tactic: Treat detector output as a weak triage signal without presenting it as provenance evidence.
- How
- Require corroborating authorship, disclosure, provenance, or editorial evidence.
- Why
- Detector accuracy is dataset- and model-dependent.
- Reference
- Synthetic Sources?
Do cited URLs resolve without errors, unsafe redirects, login walls, or removed content?
Priority: 4
- What
- Evidence-backed tactic: Measure whether users can inspect the cited evidence.
- How
- Fetch links at observation time and classify status, redirect, and access barrier.
- Why
- An inaccessible citation cannot provide practical verifiability.
Are URL parameters, fragments, redirects, and duplicate paths resolved before page-level reporting?
Priority: 4
- What
- Evidence-backed tactic: Map observed links to stable canonical identities while retaining raw URLs.
- How
- Follow safe redirects and use declared canonicals with anomaly checks.
- Why
- URL variants can inflate cited-page diversity.
Is repeated citation of the same page within one answer counted consistently?
Priority: 3
- What
- Evidence-backed tactic: Preserve events but define response-level deduplication.
- How
- Publish raw-event, unique-page, and response-prevalence metrics.
- Why
- Different interfaces repeat links differently.
Is inter-rater agreement measured for subjective GEO rubrics?
Priority: 4
- What
- Evidence-backed tactic: Quantify consistency for prominence, relevance, stance, and support judgments.
- How
- Double-score a stratified sample and adjudicate systematic disagreements.
- Why
- A human rubric is not reliable merely because humans applied it.
Are legal, safety, identity, pricing, and reputational failures confirmed by a qualified human?
Priority: 5
- What
- Evidence-backed tactic: Prevent automated scores from closing high-impact incidents.
- How
- Route critical flags to subject-matter, legal, or brand review.
- Why
- Automated graders can be confidently wrong.
- Reference
- OpenAI evaluation best practices
Is every automated LLM grader compared with blinded human judgments on representative cases?
Priority: 5
- What
- Evidence-backed tactic: Validate judge accuracy before scaling its scores.
- How
- Maintain a gold set, confusion analysis, and periodic recalibration.
- Why
- An unvalidated judge can automate bias and drift.
Are automated judges asked for classification, pairwise choice, or anchored scores rather than open-ended impressions?
Priority: 4
- What
- Evidence-backed tactic: Constrain grading to auditable decisions.
- How
- Use explicit labels, evidence requirements, and an abstain option.
- Why
- Bounded tasks are easier to calibrate and reproduce.
- Reference
- OpenAI evaluation best practices
Is every score tied to a rubric version and example set?
Priority: 4
- What
- Evidence-backed tactic: Treat criteria changes as measurement changes.
- How
- Store version, effective date, approver, and backfill policy.
- Why
- Quiet rubric drift corrupts longitudinal trends.
- Reference
- NIST AI 600-1 Generative AI Profile
Are uncertain and disputed judgments retained instead of forced into a single label?
Priority: 3
- What
- Evidence-backed tactic: Add abstain, unclear, and adjudicated states.
- How
- Report disagreement rates and store both original ratings.
- Why
- Disagreement often reveals ambiguous prompts or weak criteria.
Are Turkish answers reviewed by fluent Turkish evaluators and English answers by fluent English evaluators?
Priority: 5
- What
- Evidence-backed tactic: Evaluate meaning, fluency, cultural nuance, and factual terminology in-language.
- How
- Use native or professionally fluent reviewers with a shared cross-language rubric.
- Why
- Machine translation can hide tone, ambiguity, and factual errors.
- Reference
- NIST AI 600-1 Generative AI Profile
Are English/Turkish comparisons based on equivalent user intent rather than literal string translation?
Priority: 4
- What
- Evidence-backed tactic: Create paired prompts that are natural in each market.
- How
- Use forward translation, native rewriting, and independent intent review.
- Why
- Literal translation can change frequency, politeness, specificity, and retrieval behavior.
- Reference
- NIST AI 600-1 Generative AI Profile
Are accuracy, citation support, visibility, and incident rates disaggregated by English and Turkish?
Priority: 5
- What
- Evidence-backed tactic: Make material parity gaps visible.
- How
- Show language-level intervals and prohibit a blended score from hiding failure.
- Why
- Overall averages can conceal poor performance for one language.
- Reference
- NIST AI 600-1 Generative AI Profile
Do English and Turkish factual prompts map to equally current, authoritative first-party sources?
Priority: 4
- What
- Evidence-backed tactic: Audit missing translations and stale localized evidence.
- How
- Compare source-of-truth coverage and review dates across language pairs.
- Why
- A weaker Turkish evidence base can produce poorer citations and facts.
ANALYTICS, ATTRIBUTION & BUSINESS VALUE
Does analytics preserve and classify referral URLs carrying `utm_source=chatgpt.com`?
Priority: 5
- What
- Confirmed platform requirement: Create a documented ChatGPT referral channel rule.
- How
- Test landing-page collection, redirects, consent, and warehouse ingestion end to end.
- Why
- OpenAI automatically adds this source tag to clicked referral URLs.
- Reference
- OpenAI publishers and developers FAQ
Are referral headers and campaign parameters preserved through the landing stack where privacy policy permits?
Priority: 4
- What
- Experimental practice: Prevent avoidable source loss between click and analytics event.
- How
- Test CDN, shortener, redirect, consent, and cross-domain paths.
- Why
- Redirect and parameter loss can misclassify visits.
Does reporting avoid attributing unexplained Direct traffic growth to answer engines?
Priority: 5
- What
- Evidence-backed tactic: Keep Direct as unknown-source traffic unless corroborated.
- How
- Use tagged referrals, experiments, surveys, and assisted-journey evidence.
- Why
- Many unrelated mechanisms create Direct/none sessions.
Are internal and external redirects tested for UTM and referrer preservation?
Priority: 4
- What
- Evidence-backed tactic: Detect source loss on the exact AI-referral landing paths.
- How
- Run browser tests and inspect analytics events after every redirect hop.
- Why
- A valid tagged click can arrive as Direct after faulty routing.
Are first-user, session, and event-scoped traffic-source dimensions used for their intended questions?
Priority: 5
- What
- Confirmed platform requirement: Avoid mixing acquisition, visit, and attributed-conversion scope.
- How
- Label scope in every chart and semantic model.
- Why
- GA4 applies different traffic-source scopes and attribution behavior.
- Reference
- GA4 traffic-source dimensions
Are answer-engine referrals evaluated beyond last-click conversion?
Priority: 5
- What
- Evidence-backed tactic: Analyze new-user acquisition, return visits, branded search, and downstream conversion paths.
- How
- Use consented user/session paths and a stated attribution model.
- Why
- Research-oriented visits may influence a later channel.
- Reference
- GA4 traffic-source dimensions
Are platform impressions, clicked referrals, and inferred “dark” exposure labeled separately?
Priority: 5
- What
- Evidence-backed tactic: Attach an evidence class to every KPI.
- How
- Use `observed platform`, `observed click`, `modeled`, or `unknown` labels.
- Why
- Combining them hides major identifiability limits.
Are AI-referred landing pages analyzed by intent, content type, language, and device?
Priority: 4
- What
- Evidence-backed tactic: Diagnose which experiences convert or fail after the click.
- How
- Build cohorts from validated source tags and canonical landing URLs.
- Why
- Channel-wide averages can mask high- and low-value entry points.
- Reference
- GA4 traffic-source dimensions
Are conversion rates based on qualified sessions or users with bot and internal traffic removed?
Priority: 5
- What
- Evidence-backed tactic: Define eligible traffic and conversion events before comparison.
- How
- Publish numerator, denominator, exclusions, and confidence interval.
- Why
- Tiny or contaminated denominators create dramatic but unstable rates.
- Reference
- GA4 traffic-source dimensions
Are macro conversions, micro conversions, duplicates, tests, and invalid events classified and audited?
Priority: 5
- What
- Evidence-backed tactic: Maintain a conversion-event dictionary with business meaning and firing rules.
- How
- Test events end to end and distinguish task completion from engagement proxies.
- Why
- Duplicate or soft events can manufacture apparent channel performance.
- Reference
- GA4 traffic-source dimensions
Are revenue per session, qualified-lead rate, deal quality, or equivalent value metrics compared by source?
Priority: 5
- What
- Evidence-backed tactic: Evaluate outcome quality alongside whether a conversion occurred.
- How
- Define a business-accepted value metric and apply minimum-sample safeguards.
- Why
- Channels can produce the same conversion rate with very different value.
- Reference
- GA4 traffic-source dimensions
Are engagement signals tied to the landing page's purpose rather than a generic session-duration target?
Priority: 3
- What
- Evidence-backed tactic: Choose actions such as reading, tool use, download, contact, or task completion.
- How
- Validate event instrumentation and segment by source and intent.
- Why
- A short visit may represent success when the answer is immediately found.
- Reference
- Google AI features and your website
Are phone, meeting, proposal, retail, or other offline outcomes joined where lawful and useful?
Priority: 4
- What
- Evidence-backed tactic: Extend measurement beyond browser conversion events.
- How
- Use consented identifiers, controlled lookback rules, and documented match rates.
- Why
- High-consideration GEO influence may surface after the web session.
- Reference
- GA4 traffic-source dimensions
Are crawler requests, monitors, employees, agencies, and QA traffic excluded from human referral KPIs?
Priority: 5
- What
- Evidence-backed tactic: Maintain separate human analytics and crawler observability datasets.
- How
- Use validated IP, user-agent, network, and test identifiers without relying on one signal alone.
- Why
- GEO monitoring itself can contaminate the traffic it measures.
- Reference
- Perplexity crawler documentation
Are consent-denied and unmeasured visits acknowledged in GEO traffic reporting?
Priority: 5
- What
- Evidence-backed tactic: Separate observed analytics from total traffic claims.
- How
- Report consent rates, modeled fields, and market-specific collection differences.
- Why
- Measurement coverage can vary by jurisdiction, device, and browser.
- Reference
- GA4 traffic-source dimensions
Is the site's Referrer-Policy tested for unintended loss of useful, permitted referral context?
Priority: 3
- What
- Evidence-backed tactic: Understand what referrer data browsers send across origins.
- How
- Inspect headers on AI-referral landing pages and run cross-origin browser tests.
- Why
- Policy and link attributes can suppress the `Referer` header.
- Reference
- W3C Referrer Policy
Is answer exposure without a click reported separately from referral traffic?
Priority: 5
- What
- Evidence-backed tactic: Use platform impression or sampled-answer evidence as leading indicators only.
- How
- Maintain distinct dashboards and avoid estimated click-through without data.
- Why
- A visible citation may satisfy the user without a site visit.
Does reporting avoid calculating click-through rate from a platform report that only provides impressions or citations?
Priority: 5
- What
- Confirmed platform requirement: Require a compatible click numerator and impression denominator.
- How
- Disable synthetic CTR fields for current Google Generative AI and Bing citation reports unless the platform adds documented clicks.
- Why
- Mixing unrelated datasets creates a meaningless ratio.
Are AI-referral conversion comparisons stratified by major traffic-mix factors and accompanied by raw sample sizes?
Priority: 4
- What
- Evidence-backed tactic: Compare like with like before attributing conversion differences to the referral channel.
- How
- Predefine relevant strata such as intent, landing page, market, device, and new-versus-returning status. Report each stratum's numerator, denominator, and uncertainty, or fit an adjusted model.
- Why
- Different traffic composition can confound aggregate channel comparisons. Stratification can reduce systematic comparison error.
Do prompt, response, analytics, and CRM joins have approved retention periods and role-based access?
Priority: 5
- What
- Evidence-backed tactic: Minimize exposure of user and commercial data.
- How
- Document purpose, fields, access roles, deletion, and audit logging.
- Why
- A richer GEO dataset creates a larger privacy and security surface.
- Reference
- NIST AI 600-1 Generative AI Profile
EXPERIMENTATION
Is the “up to 40% GEO lift” claim confined to its actual experiment?
Priority: 1
- What
- Experimental practice: Describe it only as a relative maximum on the study's visibility metric after five sources were already supplied in a fixed context.
- How
- Cite the original paper with the July 2026 critical survey. State that the setup was post-retrieval and did not test stable crawling, indexing, cross-platform discoverability, traffic, or business outcomes.
- Why
- The figure is widely misreported as a general organic-visibility or click lift that the evidence does not establish.
- Reference
- Foundational GEO paper
- Reference
- Critical survey of GEO research
Are added citations, quotations, and statistics tested only when truthful and useful?
Priority: 1
- What
- Experimental practice: Evaluate whether extractable evidence improves source use without degrading retrieval or readability.
- How
- A/B test verified additions against an untreated baseline and audit factual support.
- Why
- The KDD experiment found post-retrieval gains for these tactics.
- Reference
- Foundational GEO paper
Has “authoritative tone” been rejected as a substitute for authority?
Priority: 1
- What
- Experimental practice: Do not make prose more forceful merely to influence a model.
- How
- improve evidence, expertise, qualifications, and uncertainty instead. Test tone only for reader clarity.
- Why
- The foundational paper found no significant general improvement from authoritative styling.
- Reference
- Foundational GEO paper
Has the team rejected a universal tiny-chunk or ideal-length rule?
Priority: 4
- What
- Evidence-backed tactic: Structure content around reader needs and coherent topics.
- How
- require evidence before enforcing paragraph, section, or page-length thresholds.
- Why
- Google explicitly says tiny “chunking” is not required and there is no ideal page length.
- Reference
- Google AI optimization guide
Is `llms.txt` excluded from Google visibility requirements and promises?
Priority: 4
- What
- Evidence-backed tactic: Treat it as an optional file only for consumers that explicitly document support.
- How
- remove Google ranking claims and do not divert maintenance from indexable content.
- Why
- Google says it ignores `llms.txt` for Search and generative features.
- Reference
- Google AI optimization guide
Has AI-only rewriting and exact prompt-variant publishing been rejected?
Priority: 4
- What
- Evidence-backed tactic: Write naturally for people and consolidate equivalent phrasings.
- How
- review briefs for long-tail variant pages, synonym stuffing, or a synthetic “AI voice.”
- Why
- Google says no special AI writing style or every-query variation is needed.
- Reference
- Google AI optimization guide
Are fabricated reviews, paid undisclosed mentions, and coordinated fake citations prohibited?
Priority: 5
- What
- Evidence-backed tactic: Reject inauthentic corroboration.
- How
- audit outreach, affiliate, influencer, review, and digital-PR programs for identity, disclosure, editorial independence, and spam.
- Why
- Google says seeking inauthentic mentions is unhelpful and spam systems still apply.
- Reference
- Google AI optimization guide
Has special “GEO schema” or guaranteed citation markup been rejected?
Priority: 4
- What
- Evidence-backed tactic: Use structured data for truthful semantics and documented features.
- How
- require an official consumer specification before adding a property and never promise AI inclusion.
- Why
- Google says structured data is not required for generative AI and no special schema exists.
- Reference
- Google AI optimization guide
Does every GEO test distinguish post-retrieval effects from organic discoverability?
Priority: 2
- What
- Experimental practice: Name whether the document was injected, retrieved, reranked, cited, absorbed, clicked, or converted.
- How
- prespecify the causal stage and denominator before data collection.
- Why
- Fixed-context studies cannot establish crawling or retrieval effects.
- Reference
- Critical GEO survey
Are rewrites checked for upstream retrieval harm before rollout?
Priority: 2
- What
- Experimental practice: Ensure a passage optimized for citation has not lost query relevance or retrievability.
- How
- compare index/retrieval/rerank/citation outcomes and reader quality against the original.
- Why
- The survey reports an end-to-end benchmark where body-only optimization reduced upstream and final outcomes.
- Reference
- Critical GEO survey
Are content experiments repeated across named engines and modes?
Priority: 2
- What
- Experimental practice: Avoid generalizing one answer surface to “AI search.”
- How
- record product, mode, version/date, account state, locale, and the exact prompts for each run.
- Why
- Commercial engines differ and vary over time.
- Reference
- Critical GEO survey
Do experiments include paraphrases, repeated runs, time windows, baselines, and null outputs?
Priority: 2
- What
- Experimental practice: Estimate variability rather than showcase a favorable screenshot.
- How
- use multiple natural paraphrases, close repetitions, later reruns, untreated/placebo controls, and retain no-search/no-citation/error outcomes.
- Why
- The survey recommends these controls for reproducible GEO measurement.
- Reference
- Critical GEO survey
Are citation counts paired with support, attribution, tone, and factual-use audits?
Priority: 2
- What
- Experimental practice: Measure whether the answer used the source correctly.
- How
- sample cited claims and rate entailment, accuracy, polarity, prominence, and missing qualifications with human review.
- Why
- Citation presence can coexist with unsupported or negative use.
- Reference
- Critical GEO survey
Are citations separated from clicks, conversions, and business value?
Priority: 2
- What
- Experimental practice: Treat visibility as a vector rather than one rank.
- How
- report discoverability, citation, prominence, factual absorption, referral, assisted outcome, and conversion separately.
- Why
- Microsoft says citation counts do not indicate ranking or placement. Research finds no stable downstream causal result.
- Reference
- Bing AI Performance
Does source-quality monitoring include concentration, synthetic-source risk, and diversity?
Priority: 2
- What
- Experimental practice: Audit what kinds of sources answer engines cite around priority topics.
- How
- classify ownership, originality, credibility, AI-generation evidence, geography, and repeated domain concentration. Manually verify flags.
- Why
- Recent audits find concentrated and sometimes synthetic source ecosystems.
- Reference
- Synthetic Sources audit
Does every GEO change have a prewritten hypothesis, target segment, metric, and expected direction?
Priority: 5
- What
- Evidence-backed tactic: Define what result would support or refute the change.
- How
- Register the hypothesis before publishing the treatment.
- Why
- Post-hoc stories make every outcome look successful.
- Reference
- OpenAI evaluation best practices
Is the main content or technical change distinguishable from unrelated edits?
Priority: 5
- What
- Evidence-backed tactic: Limit simultaneous variables or use a factorial design.
- How
- Keep a change manifest and postpone unrelated updates on test units.
- Why
- Bundled changes prevent causal learning.
- Reference
- NIST AI 600-1 Generative AI Profile
Does the experiment include a comparable untreated control during the same period?
Priority: 5
- What
- Evidence-backed tactic: Separate treatment effects from platform-wide movement.
- How
- Match or randomize eligible pages, topics, markets, or prompt clusters.
- Why
- Before/after change alone is confounded by time.
- Reference
- NIST AI 600-1 Generative AI Profile
Is randomization performed at a unit that avoids treatment contamination?
Priority: 4
- What
- Evidence-backed tactic: Choose page, topic cluster, locale, or market deliberately.
- How
- Document the unit, blocking variables, and spillover risks.
- Why
- Randomizing individual URLs can fail when engines synthesize across a whole site.
- Reference
- NIST AI 600-1 Generative AI Profile
Are treatment and control pages comparable in intent, baseline demand, authority, freshness, and language?
Priority: 4
- What
- Evidence-backed tactic: Reduce baseline imbalance before comparison.
- How
- Pair or stratify units using pre-period data and domain expertise.
- Why
- Page selection bias can dominate the treatment.
- Reference
- OpenAI evaluation best practices
Do users and Googlebot receive the same experiment logic and content eligibility?
Priority: 5
- What
- Evidence-backed tactic: Avoid bot-specific variants.
- How
- Use normal client/server experimentation without user-agent targeting.
- Why
- Google identifies test-page cloaking as a spam-policy violation.
- Reference
- Google A/B testing best practices
Do alternate test URLs point to the preferred original with `rel="canonical"` where appropriate?
Priority: 5
- What
- Confirmed platform requirement: Signal the preferred URL during multi-URL tests.
- How
- Validate rendered tags and ensure the canonical matches the experiment design.
- Why
- Google recommends canonical links rather than `noindex` for alternate test URLs.
- Reference
- Google A/B testing best practices
Do redirect-based temporary experiments use a temporary redirect rather than a permanent one?
Priority: 4
- What
- Confirmed platform requirement: Preserve the original URL as the intended long-term destination.
- How
- Use the documented temporary status and validate caching behavior.
- Why
- A permanent redirect sends the wrong persistence signal for a test.
- Reference
- Google A/B testing best practices
Are minimum duration, collection volume, and maximum duration fixed before the experiment starts?
Priority: 5
- What
- Evidence-backed tactic: Avoid ending on a favorable fluctuation or leaving variants indefinitely.
- How
- Base duration on crawl/recrawl latency, variance, and business risk.
- Why
- Engines may take time to discover changes and measurements are stochastic.
- Reference
- Google AI features and your website
Are campaigns, news, product launches, holidays, and demand shifts annotated and modeled?
Priority: 4
- What
- Evidence-backed tactic: Identify external events that affect both prompts and citations.
- How
- Use concurrent controls and event annotations.
- Why
- Changes in the world may affect visibility even when the page is unchanged.
For non-random tests, is treatment change compared with control change under a checked parallel-trends assumption?
Priority: 3
- What
- Experimental practice: Estimate relative change rather than treatment-only before/after change.
- How
- Plot pre-trends, disclose violations, and run sensitivity checks.
- Why
- Quasi-experiments can improve inference but rely on strong assumptions.
- Reference
- NIST AI 600-1 Generative AI Profile
Are experiments reported as treatment effects with uncertainty rather than only percent lift?
Priority: 5
- What
- Evidence-backed tactic: Show absolute and relative change, interval, sample size, and baseline.
- How
- Use the precommitted estimator and retain null or negative results.
- Why
- Percent lift alone exaggerates small baselines and hides uncertainty.
Are many prompts, platforms, metrics, and segments accounted for when declaring a win?
Priority: 4
- What
- Evidence-backed tactic: Preselect primary outcomes and control false discoveries.
- How
- Use an appropriate correction or clearly label exploratory analyses.
- Why
- Testing enough slices will produce chance “wins.”
- Reference
- OpenAI evaluation best practices
Are human graders blinded to treatment, date, and desired outcome when feasible?
Priority: 4
- What
- Evidence-backed tactic: Reduce expectancy bias in qualitative scoring.
- How
- Randomize answer order and remove treatment identifiers.
- Why
- Knowing which variant “should” win can influence ratings.
Are SEO traffic, conversions, accessibility, legal accuracy, and brand-safety guardrails monitored with rollback criteria?
Priority: 5
- What
- Evidence-backed tactic: Protect critical outcomes while testing GEO changes.
- How
- Set thresholds, owner, rollback mechanism, and evidence-preservation step before launch.
- Why
- A visibility experiment can harm users or established search performance.
- Reference
- NIST AI 600-1 Generative AI Profile
Are hypotheses, variants, dates, units, metrics, results, and decisions stored in a searchable registry?
Priority: 4
- What
- Evidence-backed tactic: Preserve institutional learning and prevent repeated tests.
- How
- Require a registry ID in deployment notes and dashboards.
- Why
- Unrecorded tests create unexplained trend breaks and selective memory.
- Reference
- OpenAI evaluation best practices
MONITORING & INCIDENT RESPONSE
Are Google and Bing AI-report exports archived with property, filters, timezone, and export date?
Priority: 4
- What
- Evidence-backed tactic: Preserve first-party observations beyond changing dashboards.
- How
- Store immutable exports and a manifest without scraping unsupported fields.
- Why
- Limited retention, rollout, and product changes can break trend reconstruction.
Are site releases, content updates, crawler-policy changes, and known platform changes overlaid on trends?
Priority: 5
- What
- Evidence-backed tactic: Maintain a shared change ledger.
- How
- Link deployments and incidents to measurement windows.
- Why
- Unannotated changes encourage false causal stories.
- Reference
- NIST AI 600-1 Generative AI Profile
Do visibility alerts account for expected variance, sample size, and repeated breaches?
Priority: 5
- What
- Evidence-backed tactic: Alert on material, sustained deviation rather than every point change.
- How
- Calibrate thresholds from baseline intervals and include data-quality gates.
- Why
- Deterministic thresholds create alert fatigue on stochastic systems.
Are GEO incidents classified by factual, legal, safety, identity, reach, and business impact with response targets?
Priority: 5
- What
- Evidence-backed tactic: Create severity levels, owners, and escalation routes.
- How
- Map critical claim types to acknowledgement, investigation, and remediation targets.
- Why
- A false address and a false medical or legal claim should not share one queue.
- Reference
- NIST AI 600-1 Generative AI Profile
Is there a documented path to correct owned facts and report persistent answer-engine errors?
Priority: 5
- What
- Evidence-backed tactic: Separate source repair, indexing request, platform feedback, and stakeholder communication.
- How
- Preserve evidence, update the authoritative page, verify recrawl, and resample before closure.
- Why
- Repeated prompting without correcting source truth is not remediation.
- Reference
- NIST AI 600-1 Generative AI Profile
Are prompt, response, citations, screenshots, source snapshots, timestamps, and context preserved for material incidents?
Priority: 5
- What
- Evidence-backed tactic: Create an auditable incident bundle.
- How
- Use controlled storage, legal retention rules, and immutable identifiers.
- Why
- Stochastic answers may disappear before investigation.
- Reference
- NIST AI 600-1 Generative AI Profile
LEGAL, BRAND & VENDOR GOVERNANCE
Are high-impact facts proactively tested with adversarial, ambiguous, and outdated formulations?
Priority: 5
- What
- Evidence-backed tactic: Probe identity, leadership, pricing, guarantees, safety, legal status, and crisis narratives.
- How
- Run a risk-ranked bilingual suite and route failures to qualified owners.
- Why
- Severe misinformation may be rare in a general benchmark.
- Reference
- NIST AI 600-1 Generative AI Profile
Are superlatives, performance claims, comparisons, guarantees, and statistics backed by current evidence and conditions?
Priority: 5
- What
- Evidence-backed tactic: Maintain claim substantiation and expiry records.
- How
- Link approved claims to evidence and remove or update stale statements across languages.
- Why
- Answer engines can repeat unsupported marketing language as fact.
- Reference
- FTC Endorsement Guides FAQ
Are paid, employee, affiliate, gifted, or otherwise material relationships disclosed clearly near endorsements?
Priority: 5
- What
- Confirmed legal requirement: Keep endorsements honest and non-misleading, and disclose relationships that could materially affect how an audience evaluates them.
- How
- Make each disclosure clear and conspicuous in its actual format and context. Test placement, prominence, wording, and repetition rather than relying on a generic notice.
- Why
- The FTC treats an undisclosed material relationship as a deception risk and provides no universal wording or placement safe harbor.
- Reference
- FTC Endorsement Guides FAQ
Has counsel assessed whether AI interactions, deepfakes, or public-interest text fall within EU AI Act Article 50 duties?
Priority: 5
- What
- Confirmed legal requirement: Treat 2 August 2026 as the general application date after a role, content, audience, jurisdiction, and exemption analysis.
- How
- Record provider/deployer status, human-review process, disclosure method, and system placement date. The grace to 2 December 2026 is limited to systems placed on the market before 2 August and the Article 50(2) machine-readable marking and detection duty.
- Why
- The duties are material and date-sensitive but contextual. They do not require a blanket label for every AI-assisted edit.
Where an EU public-interest text exemption relies on human review or editorial control, is that review substantive and documented?
Priority: 5
- What
- Confirmed legal requirement: Preserve reviewer expertise, factual checks, approval authority, and editorial responsibility.
- How
- Require content-level approval records rather than only spelling or grammar checks.
- Why
- The Commission says superficial procedural checks are not sufficient human review.
Do high-risk synthetic or materially edited assets preserve verifiable provenance through publishing transformations?
Priority: 4
- What
- Evidence-backed tactic: Carry creation, edit, ingredient, and signer context where supported.
- How
- Validate the C2PA manifest after optimization, CDN processing, and download/re-upload flows.
- Why
- Provenance can help users and systems understand how an asset changed.
Does policy treat Content Credentials as provenance evidence rather than proof that an asset or claim is true?
Priority: 5
- What
- Evidence-backed tactic: Use a validated credential to assess signed provenance and asset integrity. A credential is not a truth score.
- How
- Review the signer, validation state, assertions, and content binding, then fact-check the depicted or stated claim independently.
- Why
- C2PA says valid manifests can coexist with misinformation. They establish verifiable association and freedom from tampering without establishing truth.
- Reference
- C2PA 2.4 harms modelling
Are rights, licenses, attribution, quotation limits, and AI-use restrictions checked before publishing source-derived GEO content?
Priority: 5
- What
- Evidence-backed tactic: Maintain a rights record for text, images, datasets, testimonials, and generated assets.
- How
- Route uncertain reuse and training-origin questions to qualified counsel.
- Why
- Search visibility does not grant permission to copy or republish.
- Reference
- US Copyright Office AI initiative
Are GEO vendors assessed for methodology, data rights, retention, security, model changes, exports, and incident duties?
Priority: 4
- What
- Evidence-backed tactic: Avoid outsourcing accountability to an opaque visibility score.
- How
- Contract for definitions, raw evidence access, change notice, deletion, and exit portability.
- Why
- Vendor methodology changes can silently rewrite historical results.
- Reference
- NIST AI 600-1 Generative AI Profile
FRONTIER RIGHTS PROTOCOLS: EXPERIMENTAL
Where machine-readable AI licensing is part of the rights strategy, is every covered asset associated with a valid RSL 1.0 license?
Priority: 1
- What
- Experimental practice: Publish explicit permissions, prohibitions, licensing, attribution, or payment terms for automated uses.
- How
- Create an `application/rsl+xml` document in the RSL 1.0 namespace, scope its `<content>` rules carefully, and expose it through a supported discovery mechanism.
- Why
- RSL 1.0 defines an industry-recommendation format that distinguishes search, AI indexing and input, training, attribution, and payment terms.
- Reference
- RSL 1.0 specification
Are RSL 1.0 associations validated for media type, scope, precedence, and cross-channel consistency?
Priority: 2
- What
- Experimental practice: Make the effective license unambiguous for every covered asset.
- How
- Test `application/rsl+xml`, absolute references, user-agent and asset scope, specific-over-broad precedence, and conflicts across robots.txt, HTTP `Link`, HTML, RSS, and embedded metadata.
- Why
- RSL requires clients to evaluate discoverable associations and applies specificity and restrictive conflict-resolution rules.
If Content Signals are used, do `search`, `ai-input`, and `ai-train` reflect separately approved choices rather than inherited assumptions?
Priority: 1
- What
- Experimental practice: Express post-access use preferences separately from crawler access rules.
- How
- Publish the Content Signals Policy and only approved `Content-Signal` values in robots.txt. Leave an undecided use unspecified and test the effective file after CDN composition.
- Why
- Cloudflare defines separate signals for search indexing, real-time AI input, and training. If a signal is absent, this mechanism neither grants nor restricts anything.
Where counsel selects TDMRep for applicable rights reservations, are the effective `tdm-reservation` and optional `tdm-policy` values valid and consistent?
Priority: 2
- What
- Experimental practice: Expose a machine-readable reservation of text-and-data-mining rights and, when offered, a route to licensing terms.
- How
- Publish `/.well-known/tdmrep.json` or supported HTTP, HTML, or asset metadata. Use `tdm-reservation: 1` to reserve rights, link a policy where applicable, and test precedence.
- Why
- The TDMRep Final Community Group Report defines interoperable reservation and policy discovery for lawfully accessible web content.
PLATFORM REPORTING & DIAGNOSTICS: VERIFIED 23 JULY 2026
Is every property verified in Google Search Console?
Priority: 5
- What
- Confirmed platform requirement: Establish official crawl/index diagnostics for Google AI eligibility.
- How
- verify all protocols/subdomains or a suitable domain property and assign least-privilege access.
- Why
- Google directs site owners to Search Console for AI-feature technical diagnosis.
- Reference
- Google AI features and websites
Does Google URL Inspection show the intended fetched HTML, canonical, index state, and preview controls?
Priority: 5
- What
- Confirmed platform requirement: Observe what Google received.
- How
- inspect representative pages after every template or edge-policy release.
- Why
- Google recommends it when AI preview controls appear ineffective.
- Reference
- Google AI features and websites
Are Google generative-AI impressions reconciled with overall Web performance and analytics conversions?
Priority: 4
- What
- Confirmed platform requirement: Avoid double-counting or mistaking a rollout gap for zero visibility.
- How
- document report availability, compare URL/country/time trends, and join landing outcomes carefully.
- Why
- Dedicated data remains included in overall performance.
Are Bing Site Explorer error, warning, excluded, noindex, robots, redirect, and malware cohorts monitored?
Priority: 4
- What
- Confirmed platform requirement: Detect sitewide retrieval regressions.
- How
- baseline counts by folder and alert on discontinuities.
- Why
- Bing exposes these crawl/index cohorts directly.
- Reference
- Bing Site Explorer
Are crawler logs retained with verified vendor identity, status, bytes, latency, cache result, and route?
Priority: 5
- What
- Evidence-backed tactic: Build ground truth for access incidents.
- How
- enrich edge/origin logs after IP/reverse-DNS verification and chart by product token.
- Why
- UA strings alone are spoofable and platform consoles are incomplete.
- Reference
- Googlebot documentation
Does a scheduled synthetic matrix fetch key URLs as each supported crawler from relevant regions?
Priority: 5
- What
- Evidence-backed tactic: Catch WAF, CDN, locale, and response drift before visibility falls.
- How
- test robots plus representative HTML/assets using documented UA/IP-safe methods.
- Why
- Intended policy can diverge from effective transport.
- Reference
- Perplexity crawler documentation
Are crawler-specific `403`, `401`, `429`, `5xx`, challenge, and zero-byte spikes alerted quickly?
Priority: 5
- What
- Evidence-backed tactic: Detect access regressions by bot and layer.
- How
- set anomaly alerts and preserve sampled request/response traces.
- Why
- These failures block retrieval or throttle crawling.
- Reference
- OpenAI advertiser crawler guidance
Is every robots/WAF change revalidated after each platform's documented propagation interval?
Priority: 4
- What
- Evidence-backed tactic: Close the deployment loop.
- How
- capture before/after parser output, live requests, logs, and platform diagnostics at the right delay.
- Why
- OpenAI and Perplexity document up-to-about-24-hour adjustment periods. Google requires recrawl.
- Reference
- OpenAI crawler overview
Is a stable, versioned prompt panel used to observe answer-surface eligibility and citation changes?
Priority: 3
- What
- Experimental practice: Measure real outputs across engines, locales, devices, and fresh/anonymous sessions.
- How
- freeze prompt intent, record date/model/surface/citations, and repeat without treating outcomes as rank truth.
- Why
- No publisher console covers all answer engines.
- Reference
- ChatGPT Search help
Does the stale-content incident runbook update/delete the source, transport the change, and verify removal on every relevant index?
Priority: 5
- What
- Evidence-backed tactic: Handle harmful outdated answers end to end.
- How
- correct or remove the page, return the right status/directive, update sitemap/IndexNow/feed, request recrawl, and use platform removal channels if necessary.
- Why
- No single robots change clears every cached/indexed surface.
- Reference
- Anthropic blocking and removal
Are platform documentation and crawler-token changes reviewed on a fixed cadence?
Priority: 4
- What
- Evidence-backed tactic: Detect renamed bots, new controls, changed IP endpoints, and reporting features.
- How
- diff official pages monthly and trigger policy review on semantic changes.
- Why
- Multiple reviewed pages changed in July 2026 alone.
- Reference
- Google common crawlers
Are Google property-level and page-filtered Generative AI impression totals interpreted using their different aggregation rules?
Priority: 4
- What
- Confirmed platform requirement: Preserve report scope and aggregation mode with exports.
- How
- Do not reconcile totals by simple summation across pages.
- Why
- Multiple results from one property can count differently after URL filtering.
Does the Google Generative AI report avoid attributing impressions to prompts or grounding queries it does not expose?
Priority: 5
- What
- Confirmed platform requirement: Use only documented dimensions: pages, countries, dates, and devices.
- How
- Label private prompt sampling as a separate dataset.
- Why
- Joining unrelated prompt tests to platform impressions creates false precision.
Are Bing grounding queries described as examples rather than a complete demand or keyword dataset?
Priority: 5
- What
- Confirmed platform requirement: Preserve Bing's sampling caveat in downstream reports.
- How
- Prohibit extrapolation to search volume or total prompt share.
- Why
- Sampled grounding queries do not define the full retrieval universe.
Is a Google Generative AI impression counted only according to the platform's displayed-link definition?
Priority: 5
- What
- Confirmed platform requirement: Keep first-party report semantics distinct from private answer mentions.
- How
- Copy the current definition into the metric dictionary and date it.
- Why
- Private “visibility” observations are not interchangeable with Search Console impressions.
Is Google Generative AI report availability recorded rather than treating an absent report as zero visibility?
Priority: 4
- What
- Confirmed platform requirement: Distinguish unavailable, insufficient-data, and zero-impression states.
- How
- Add an availability flag and screenshot the report status.
- Why
- The report is not available to every property.
Are the newest Google Generative AI data points marked provisional until stabilized?
Priority: 3
- What
- Confirmed platform requirement: Prevent premature incident or success conclusions.
- How
- Reconcile recent periods after the documented processing window.
- Why
- Google says newest data can be preliminary and later change.
You've reached the end of 449 checks.
Every item shows its source, and the page carries the date it was last verified. That date does more work here than it would elsewhere. Crawler names, access controls and reporting all get revised without much warning, so a check that was accurate then can quietly stop being accurate now. None of it guarantees a citation. If classic search is the bigger gap right now, the SEO checklist is the other half of this.
Tell us what you're fixing