Two specialists use magnifying glasses to audit a GEO checklist and linked sources

Verified 23 July 2026: audit 449 controls across 30 domains for AI-search eligibility, evidence quality, observability, and operational readiness. Every check includes What, How, Why, and a direct source, labeled as a confirmed platform requirement, confirmed legal requirement, evidence-backed tactic, or experimental practice. Priority 5 is critical. Priority 1 is experimental or optional. No item guarantees ranking, visibility, traffic, or citation.

Is there a current crawler-to-purpose inventory?

Priority: 5
What
Confirmed platform requirement: Classify every platform behavior as automatic crawl, training crawl, user-triggered fetch, or answer-time reuse or grounding of indexed content.
How
Record each product token, user-agent pattern, control mode, owner, intended paths, and official source. Do not assume one rule governs the other modes.
Why
The same vendor can expose independent controls for access, training, user actions, and answer context.

Has the organization approved separate retrieval and training policies?

Priority: 5
What
Evidence-backed tactic: Decide search visibility independently from foundation-model training.
How
Maintain a signed policy matrix for Googlebot/Google-Extended, OAI-SearchBot/GPTBot, Claude-SearchBot/ClaudeBot, and Applebot/Applebot-Extended.
Why
Blocking training need not sacrifice search.

Does every crawler-policy change have an accountable owner and review date?

Priority: 5
What
Evidence-backed tactic: Assign legal, security, SEO, and engineering responsibility.
How
Store intent, approver, effective date, and next review beside the rule.
Why
Fast-changing bot names and uses make undocumented rules unsafe.

Are robots.txt rules version-controlled and tested before release?

Priority: 5
What
Evidence-backed tactic: Treat crawler policy as production configuration.
How
Diff parsed groups, test representative URLs per token, and retain rollback evidence.
Why
A malformed or overbroad rule can remove an entire answer surface.

Is the policy deployed on every relevant hostname and subdomain?

Priority: 5
What
Confirmed platform requirement: Cover each separately hosted locale, media host, CDN, docs site, and subdomain.
How
Fetch `/robots.txt` and target URLs on each host.
Why
Robots rules are host-scoped.

Are automatic crawlers distinguished from user-triggered fetchers?

Priority: 5
What
Evidence-backed tactic: Document whether user fetches obey the same robots policy.
How
Test and record ChatGPT-User, Claude-User, Perplexity-User, and platform-specific semantics.
Why
A search-crawl allow or block may not govern a user request.

Are official IP-range endpoints treated as changing dependencies?

Priority: 5
What
Evidence-backed tactic: Refresh allowlists rather than copying static IPs into policy forever.
How
Pull vendor JSON endpoints on a controlled schedule and alert on changes.
Why
Stale allowlists silently block legitimate bots.

Do stakeholder claims avoid promising inclusion, rank, or citation?

Priority: 5
What
Confirmed platform requirement: Describe controls as eligibility and discovery measures.
How
Remove guaranteed-visibility language from briefs and dashboards.
Why
Platforms reserve selection decisions.

Is there a documented removal and escalation path per answer engine?

Priority: 4
What
Evidence-backed tactic: Cover emergency de-indexing, stale citations, copyright, and crawler malfunction.
How
Record platform tools, evidence required, contacts, and validation steps.
Why
Robots changes may not remove already indexed material immediately.

Are third-party search-index dependencies included in incident analysis?

Priority: 4
What
Evidence-backed tactic: Recognize that an answer engine may use partner indexes.
How
Check both the answer engine and underlying search-engine access/index state.
Why
Allowing one proprietary bot may not restore partner-fed results.

Does a recurring review compare English and Turkish visibility, accuracy, citations, referrals, experiments, incidents, and unresolved risks?

Priority: 5
What
Evidence-backed tactic: Give owners one balanced scorecard without collapsing unlike metrics.
How
Review platform-specific panels, uncertainty, language gaps, incidents, and next decisions with local owners.
Why
A complete GEO program joins measurement, operations, business outcomes, and governance across both audiences.

Is every priority content asset tied to a researched user need?

Priority: 5
What
Evidence-backed tactic: Define the decision, task, or understanding a real audience needs.
How
Attach interview, support, sales, search, or product evidence to a concise need statement.
Why
Content built around an internal keyword list can miss the user's actual problem.

Does the page have one explicit primary audience and purpose?

Priority: 5
What
Evidence-backed tactic: State whom the content serves and what successful use looks like.
How
Put audience, context, and intended outcome in the brief and test the finished page against them.
Why
A page that tries to serve incompatible audiences becomes ambiguous.

Are content opportunities grouped by information need rather than exact keyword wording?

Priority: 4
What
Evidence-backed tactic: Consolidate synonymous prompts that seek the same answer.
How
Map variants to one canonical need and split only when intent or required evidence differs.
Why
Google says AI systems understand synonyms and meanings, so variant-page proliferation adds little value.

Does the research cover informational, comparative, transactional, and action-oriented intents?

Priority: 4
What
Evidence-backed tactic: Represent the different jobs users ask an answer engine to perform.
How
Classify prompt evidence by intent and document the content or tool that resolves each class.
Why
A purely informational library misses decisions and next actions.

Are likely follow-up questions mapped as a coherent journey?

Priority: 4
What
Evidence-backed tactic: Capture what a user needs before and after the initial answer.
How
Analyze support threads, conversational sessions, and related questions. Connect each stage with contextual links.
Why
Generative search is conversational and may issue additional targeted searches.

Are complex questions decomposed into genuine subproblems?

Priority: 4
What
Evidence-backed tactic: Identify definitions, eligibility, options, evidence, risks, steps, and outcomes required for a complete answer.
How
Build a topic map and assign each subproblem to a section or linked specialist page.
Why
Google AI features can use query fan-out across subtopics and data sources.

Does the content answer meaningful qualifiers and constraints?

Priority: 4
What
Evidence-backed tactic: Cover audience, budget, timing, geography, compatibility, eligibility, and risk qualifiers that change the answer.
How
Derive qualifiers from real decisions and show how recommendations change.
Why
ChatGPT may rewrite a prompt using location or remembered preferences.

Are comparison questions organized around explicit decision criteria?

Priority: 4
What
Evidence-backed tactic: Name the dimensions that distinguish options.
How
Define audience, criteria, weights or tradeoffs, evidence date, and cases where each choice fits.
Why
Comparison content is more useful when readers can audit the reasoning.

Are “when not to,” exclusion, and no-answer needs represented?

Priority: 3
What
Evidence-backed tactic: Identify situations where the product, method, or recommendation does not apply.
How
Add contraindications, prerequisites, unsupported cases, and escalation paths.
Why
Completeness includes boundaries alongside persuasive positive claims.

Are actual grounding queries used to revise topic coverage?

Priority: 4
What
Evidence-backed tactic: Treat retrieval phrases reported by Microsoft as evidence of how systems map content.
How
Review Bing Webmaster Tools or Clarity grounding-query samples, validate user relevance, then fill material gaps.
Why
Grounding queries may differ from user wording.

Do customer-facing teams contribute real questions and terminology?

Priority: 4
What
Evidence-backed tactic: Use language from support, sales, onboarding, research, and community interactions.
How
Maintain a consent-safe question repository with frequency, audience, resolution, and last-seen date.
Why
Real questions reveal needs and vocabulary that search tools miss.

Are emerging questions distinguished from short-lived trend chasing?

Priority: 3
What
Evidence-backed tactic: Separate durable needs from news-driven spikes.
How
require audience fit, evidence availability, maintenance owner, and expected shelf life before publishing.
Why
Google warns against writing merely because a topic is trending.

Does each topic cluster have a clear canonical answer and specialist depth?

Priority: 3
What
Evidence-backed tactic: Prevent multiple pages from giving inconsistent partial answers.
How
name the canonical overview, link to detailed evidence, and retire or consolidate redundant pages.
Why
Coherent internal relationships help users traverse fan-out subtopics.

Is a cross-engine, cross-language prompt panel maintained for priority needs?

Priority: 2
What
Experimental practice: Observe how named answer surfaces phrase and decompose the same need.
How
test natural prompts, several paraphrases, languages, locations, and dates. Retain outputs and null results.
Why
engines and contexts retrieve different source ecosystems.

Does each answer make the user's safe next step explicit?

Priority: 4
What
Evidence-backed tactic: Close the information gap with a suitable action, decision aid, tool, source, or escalation.
How
test that the next step follows from the evidence and works without hidden prerequisites.
Why
People-first content should leave users feeling they learned enough to achieve their goal.

Does the title accurately summarize the page's distinctive answer?

Priority: 5
What
Evidence-backed tactic: Use a descriptive, non-sensational title tied to the real content.
How
compare the title with the primary need, scope, and conclusion. Remove unsupported superlatives.
Why
Users and retrieval systems need a reliable statement of page purpose.

Is the primary answer or conclusion available before extended detail?

Priority: 4
What
Evidence-backed tactic: Give a concise, qualified answer early.
How
state the result, audience, material caveat, and evidence date, then expand.
Why
Readers can orient quickly while deeper context remains available.

Do section headings describe their actual topic or purpose?

Priority: 5
What
Evidence-backed tactic: Make headings predictive rather than clever or generic.
How
read the heading outline alone and verify it communicates the page's logic.
Why
Descriptive headings help people navigate and expose content relationships.

Is the visible hierarchy encoded with real semantic headings, lists, and tables?

Priority: 5
What
Evidence-backed tactic: Preserve relationships in markup rather than relying on styling alone.
How
inspect the DOM and accessibility tree for heading levels, list semantics, table headers, and regions.
Why
Programmatic relationships survive alternate presentations and agent parsing.

Does each section resolve one coherent subquestion without arbitrary micro-fragmentation?

Priority: 4
What
Evidence-backed tactic: Use sections where topic boundaries naturally change.
How
merge tiny fragments that lack context and split only when a reader gains navigation or comparison value.
Why
Google explicitly says tiny content “chunks” are not required.

Can important paragraphs be understood with their necessary qualifiers intact?

Priority: 4
What
Evidence-backed tactic: Keep the claim, subject, condition, date, and source close together.
How
replace orphaned pronouns and detached caveats where ambiguity results.
Why
Reuse is safer when context is not stranded elsewhere.

Are central terms defined at first meaningful use?

Priority: 4
What
Evidence-backed tactic: State what a specialized term means in this context.
How
use a concise definition, expand acronyms, and link to a maintained glossary when needed.
Why
Explicit definitions reduce entity and concept ambiguity.

Are procedures presented as ordered, testable steps?

Priority: 4
What
Evidence-backed tactic: Distinguish prerequisites, actions, expected results, and failure paths.
How
use an ordered list with commands or UI labels that match the current product.
Why
Complete steps are easier to follow and verify than narrative fragments.

Do comparison tables have explicit dimensions, units, and accessible headers?

Priority: 4
What
Evidence-backed tactic: Encode row/column relationships and the basis of comparison.
How
use table captions, `th`, `scope` or equivalent associations. Note date and missing data.
Why
Microsoft recommends tables, while WCAG requires relationships to be programmatically determinable.

Are recommendations paired with tradeoffs and selection conditions?

Priority: 4
What
Evidence-backed tactic: Explain benefits, costs, risks, and who should choose differently.
How
add a decision rule and evidence behind each material recommendation.
Why
Unsupported universal advice is easy to quote incorrectly.

Are prerequisites and dependencies visible before the recommendation or procedure?

Priority: 4
What
Evidence-backed tactic: State required access, data, skills, versions, jurisdiction, and budget.
How
add a prerequisites block and link each dependency to authoritative instructions.
Why
Missing preconditions make an otherwise correct answer unusable.

Are limitations, edge cases, and unsupported scenarios explicit?

Priority: 4
What
Evidence-backed tactic: Bound the claim or method.
How
list known failure modes, exceptions, uncertainty, and escalation criteria beside the affected guidance.
Why
Caveats separated into a distant disclaimer are easily lost.

Are FAQ sections limited to real, distinct questions?

Priority: 3
What
Evidence-backed tactic: Include questions evidenced by users that add material information.
How
remove keyword-variant questions and merge answers that resolve the same need.
Why
Microsoft recommends FAQ sections, but Google rejects variant content written just for AI.

Are acronyms, symbols, and jargon expanded for the intended audience?

Priority: 3
What
Evidence-backed tactic: Make specialized language interpretable without erasing necessary precision.
How
define at first use and keep a controlled terminology list for repeated concepts.
Why
Clear language helps people process information quickly.

Are visible publication, review, and material-update signals unambiguous?

Priority: 4
What
Evidence-backed tactic: Distinguish when the page was published, reviewed, and substantively changed.
How
label dates, explain major changes, and align visible dates with structured values.
Why
Users need to judge currency without false freshness.

Has the page's set of externally verifiable claims been inventoried?

Priority: 5
What
Evidence-backed tactic: Identify facts, numbers, comparisons, legal statements, and causal assertions needing support.
How
run a claim-level editorial pass and assign a source, owner, and review date.
Why
Page-level bibliographies cannot reveal unsupported individual claims.

Does each material claim use the strongest available primary source?

Priority: 5
What
Evidence-backed tactic: Prefer laws, standards, original research, official datasets, vendor documentation, or first-hand records.
How
trace secondary summaries back to the originating evidence and cite that page.
Why
Primary sources reduce distortion and make verification easier.

Do citations link to the deepest stable page that supports the claim?

Priority: 4
What
Evidence-backed tactic: Avoid home pages, search results, and generic documentation indexes.
How
link the exact report, section, dataset, specification, or release note. Preserve DOI or stable identifier where available.
Why
Direct sources reduce verification cost.

Has a reviewer confirmed that every cited source actually supports the adjacent claim?

Priority: 5
What
Evidence-backed tactic: Test entailment, scope, and context rather than link presence.
How
open the source, locate the supporting passage, and record mismatches or missing qualifiers.
Why
AI systems can cite a page that does not support the generated statement.

Are answer sentences decomposed into independently verifiable claims before citation scoring?

Priority: 5
What
Evidence-backed tactic: Identify each factual proposition and its associated citations.
How
Use a documented claim-splitting rule and human review for complex sentences.
Why
A 2023 snapshot of four generative search engines found only 51.5% of generated sentences were fully supported. Treat this as historical evidence of the risk rather than a current platform benchmark.

Are all externally verifiable answer claims checked for citation support?

Priority: 5
What
Evidence-backed tactic: Measure the share of claims fully supported by cited evidence.
How
Review each claim against its cited sources and record partial support separately.
Why
Fluent answers can contain uncited or only partly supported factual claims.

Does every citation actually support the claim it is attached to?

Priority: 5
What
Evidence-backed tactic: Measure the share of citations that entail or substantiate the associated claim.
How
Open the cited source, inspect the relevant passage, and score full, partial, or no support.
Why
In the same historical four-engine snapshot, 74.5% of citations fully supported the associated sentence. A real citation can still be irrelevant or misused.

Are citations placed next to the claims they support?

Priority: 4
What
Evidence-backed tactic: Minimize ambiguity between evidence and assertion.
How
add inline links or markers at sentence or paragraph level and keep a readable source list for full metadata.
Why
Claim-level proximity makes both human and automated attribution easier to audit.

Are source publication date, version, and access date recorded where currency matters?

Priority: 5
What
Evidence-backed tactic: Make the evidence state reproducible.
How
show source date/version in notes and record the audit access date internally.
Why
Vendor behavior, law, prices, and research versions change.

Does every statistic state its denominator, unit, period, and geography?

Priority: 5
What
Evidence-backed tactic: Preserve the minimum context needed to interpret a number.
How
attach sample/base, numerator and denominator, units, dates, location, and source beside the statistic.
Why
Context-free numbers are easy to reuse misleadingly.

Are study method, sample, population, and exclusions summarized with research claims?

Priority: 5
What
Evidence-backed tactic: Expose how the result was produced and whom it applies to.
How
add a method note and link to the full protocol or paper.
Why
A headline effect size without design context encourages false generalization.

Do comparisons use the same baseline, unit, and observation window?

Priority: 4
What
Evidence-backed tactic: Make unlike measures visibly incomparable.
How
normalize where legitimate. Otherwise, label basis differences and avoid a synthetic rank.
Why
Comparison tables can manufacture certainty from mismatched inputs.

Are correlation, prediction, and causation distinguished?

Priority: 5
What
Evidence-backed tactic: Match claim language to the research design.
How
reserve causal verbs for randomized or defensible quasi-experimental evidence and explain confounding elsewhere.
Why
AI summaries can amplify an overclaimed causal sentence.

Are uncertainty, ranges, and material disagreement reported?

Priority: 4
What
Evidence-backed tactic: Avoid presenting estimates or contested conclusions as settled facts.
How
include confidence intervals, plausible ranges, dissenting evidence, and what would change the conclusion.
Why
Factual reliability includes uncertainty.

Are quotations exact, bounded, and attributed to their original speaker and work?

Priority: 4
What
Evidence-backed tactic: Preserve wording without laundering a paraphrase as a quote.
How
verify the source, mark omissions or additions, link it, and name speaker, work, and date.
Why
The foundational GEO study's quotation result does not justify invented or decontextualized quotes.

Does the page represent credible conflicting sources rather than cherry-pick?

Priority: 4
What
Evidence-backed tactic: Surface material evidence that would change a reader's decision.
How
document source-selection criteria, summarize the disagreement, and explain the chosen conclusion.
Why
One-sided sourcing weakens trust and portability.

Are derived calculations reproducible?

Priority: 4
What
Evidence-backed tactic: Show formula, inputs, rounding, assumptions, and source data.
How
publish a worked example or downloadable notebook/spreadsheet for material calculations.
Why
Readers can verify the result and update it when inputs change.

Are original datasets available in documented, usable formats when disclosure is safe?

Priority: 4
What
Evidence-backed tactic: Publish the evidence behind original quantitative claims.
How
provide stable downloads or API access with schema, version, creator, license, coverage, and contact.
Why
Original research becomes verifiable and independently reusable.

Is dataset provenance explicit when data are copied, transformed, or aggregated?

Priority: 4
What
Evidence-backed tactic: Distinguish republication from derivation.
How
identify the canonical original with `sameAs` for an unchanged copy and `isBasedOn` for significant transformation or aggregation.
Why
Users need to trace data lineage.

Are source licenses and reuse rights recorded?

Priority: 4
What
Evidence-backed tactic: Make legal reuse and attribution conditions visible for data, images, quotes, and code.
How
link the license, name creator/rightsholder, and retain required notices.
Why
Citation does not itself grant permission to reuse.

Are broken, redirected, retracted, and materially changed citations monitored?

Priority: 4
What
Evidence-backed tactic: Keep the evidence graph valid over time.
How
run link checks, preserve DOI/archive identifiers, review redirects, and flag retractions or changed conclusions.
Why
A live URL can still point to evidence that no longer supports the claim.

Are volatile claims assigned a review cadence and owner?

Priority: 5
What
Evidence-backed tactic: Identify prices, availability, policies, laws, product behavior, and “current” comparisons.
How
set source-specific expiry dates, alerts, and an update-or-remove workflow.
Why
Microsoft says accurate, up-to-date content matters for AI inclusion and citation.

Is there a visible correction mechanism and material-change history?

Priority: 4
What
Evidence-backed tactic: Let readers report errors and see significant corrections.
How
publish contact/reporting paths, correction date, original error, corrected statement, and affected evidence.
Why
Traceable corrections improve accountability.

Does the page contribute a genuinely distinctive viewpoint?

Priority: 5
What
Evidence-backed tactic: Add analysis or experience unavailable in commodity summaries.
How
identify the page's unique claim, evidence, or framework in the brief and verify it survives editing.
Why
Google explicitly recommends a unique point of view for generative AI search.

Is claimed first-hand experience demonstrated with verifiable detail?

Priority: 5
What
Evidence-backed tactic: Show what was used, visited, tested, observed, or implemented.
How
include dates, setup, constraints, artifacts, and outcomes while protecting sensitive data.
Why
Google contrasts first-hand review with a restatement of existing content.

Does original research publish its question, method, data basis, and limitations?

Priority: 5
What
Evidence-backed tactic: Make proprietary findings inspectable.
How
link a methodology page and expose enough data or aggregates for verification.
Why
Original reporting is more valuable when reproducible and bounded.

Does sourced content add substantial analysis rather than paraphrase?

Priority: 5
What
Evidence-backed tactic: Go beyond rewriting another page.
How
compare the draft with cited sources and require a new synthesis, application, critique, dataset, or example.
Why
Google explicitly warns against copying or rewriting without added value.

Does the page provide material value beyond existing strong sources?

Priority: 4
What
Evidence-backed tactic: Identify the unresolved gap it closes.
How
benchmark leading primary and high-quality secondary sources, then document the incremental contribution.
Why
Commodity pages add little reason to retrieve or cite another source.

Do case studies disclose starting state, intervention, time window, and outcome?

Priority: 4
What
Evidence-backed tactic: Turn anecdotes into bounded evidence.
How
report baseline, context, steps, measurement, result, confounders, and what may not generalize.
Why
Decontextualized success stories invite false causal inference.

Are failures, negative results, and tradeoffs retained?

Priority: 4
What
Evidence-backed tactic: Publish what did not work where it changes the decision.
How
record failed approaches, conditions, costs, and lessons alongside successes.
Why
Negative evidence is distinctive and reduces survivorship bias.

Can product tests or experiments be reproduced at the stated date and version?

Priority: 4
What
Evidence-backed tactic: Preserve setup, inputs, environment, and expected result.
How
publish test protocol, screenshots or logs, version numbers, and last rerun date.
Why
Product behavior drifts and unsupported “current” tests age quickly.

Are proprietary data claims accompanied by safe evidence artifacts?

Priority: 4
What
Evidence-backed tactic: Make internally derived insights auditable without leaking protected data.
How
publish aggregates, sampling rules, anonymization, schema, and a contact for access questions.
Why
Unsupported proprietary numbers are difficult to trust or reuse.

Are expert interviews linked to the interviewee's identity and original context?

Priority: 4
What
Evidence-backed tactic: Preserve who said what, when, and in what capacity.
How
obtain consent, publish transcript or relevant excerpt, link a profile, and separate opinion from fact.
Why
Attributable primary testimony is stronger than anonymous authority language.

Does the page use original examples that actually exercise the guidance?

Priority: 4
What
Evidence-backed tactic: Demonstrate the concept in a realistic scenario.
How
create examples from tested work, state assumptions, and verify all outputs.
Why
Original examples add practical information beyond definitions.

Does the page expose a reusable decision framework rather than a flat tips list?

Priority: 4
What
Evidence-backed tactic: Connect evidence, criteria, options, and outcomes.
How
publish a flow, rubric, calculator, or worked decision with limitations.
Why
A framework helps users apply expertise to their own context.

Is scaled, low-value derivative production prohibited?

Priority: 5
What
Confirmed platform requirement: Prevent mass-generated pages that add no value.
How
audit templates, automation, outsourcing, and network publishing for originality and human review.
Why
Google classifies scaled unoriginal content created to manipulate rankings as spam.

Are update dates changed only after substantive review or modification?

Priority: 5
What
Evidence-backed tactic: Reject simulated freshness.
How
define what qualifies as a material update and retain review/change evidence.
Why
Google warns against changing dates merely to make pages appear fresh.

Is the reason for publishing primarily to help the intended audience?

Priority: 5
What
Evidence-backed tactic: Check that the page serves a real need beyond anticipated AI traffic.
How
require a user outcome, evidence gap, owner, and maintenance plan in the brief.
Why
Google asks whether content has a people-first purpose.

Does each expert or editorial page show a truthful visible byline?

Priority: 5
What
Evidence-backed tactic: Identify the accountable creator rather than a vague content team when a person is responsible.
How
place the byline near the title and link it to a maintained profile.
Why
Google asks who created content and recommends accurate authorship.

Does each author profile explain relevant experience and role?

Priority: 4
What
Evidence-backed tactic: Show why the creator is qualified for this topic.
How
list verifiable credentials, first-hand experience, affiliations, disclosures, selected work, and contact path.
Why
Expertise should be demonstrable rather than implied by tone.

Is the author's expertise matched to the page's subject and risk?

Priority: 5
What
Evidence-backed tactic: Avoid using a generic expert label across unrelated domains.
How
define topic scopes and require appropriate creator or reviewer qualifications for each risk class.
Why
Google asks whether an expert or enthusiast demonstrably knows the topic.

Are expert reviewers named with their exact review scope and date?

Priority: 5
What
Evidence-backed tactic: Distinguish writing, fact-checking, medical/legal review, and editorial approval.
How
display reviewer identity, qualification, reviewed sections, review date, and change trigger.
Why
A decorative “expert reviewed” badge can mislead.

Does every byline resolve to one stable canonical author URL?

Priority: 4
What
Evidence-backed tactic: Give each creator a durable identity page.
How
use the same profile URL in visible links and Article `author.url`. Redirect retired URLs.
Why
Google says `author.url` can uniquely identify the author.

Are all visible authors represented separately in Article markup?

Priority: 4
What
Confirmed platform requirement: Preserve multi-author accountability.
How
create one Person or Organization object per author and do not merge names into one field.
Why
Google explicitly instructs publishers to include every author in markup.

Does the site clearly identify its publication and publisher?

Priority: 5
What
Evidence-backed tactic: Explain the site's mission, editorial remit, publisher, and relationship to the organization.
How
maintain an About page and link it from content and navigation.
Why
Google cites site and publisher background as a trust cue.

For news content, are ownership, company, or network relationships transparent?

Priority: 5
What
Confirmed platform requirement: Identify the entity behind the publication.
How
disclose parent company, controlling organization, and material network relationships on an accessible page.
Why
Google News transparency policy asks for this information.

For news content, is usable contact information available?

Priority: 5
What
Confirmed platform requirement: Give readers a route to the publisher.
How
provide monitored editorial and correction contacts and, where appropriate, organization contact details.
Why
Google News transparency policy calls for contact information.

Is the editorial and fact-checking process published?

Priority: 4
What
Evidence-backed tactic: Explain source standards, review levels, update rules, and how independence is protected.
How
publish a concise policy and link it from high-stakes content.
Why
Readers can assess how claims reached publication.

Is there a public corrections policy with a working reporting path?

Priority: 5
What
Evidence-backed tactic: Define what gets corrected, how quickly, and how changes are shown.
How
test the reporting channel and sample recent corrections.
Why
Google requires a corrections policy or error-reporting mechanism for fact-check eligibility.

Are sponsorship, payment, affiliate interest, and material support clearly disclosed?

Priority: 5
What
Confirmed platform requirement: Separate commercial influence from independent editorial content.
How
label sponsorship near the affected content and explain the relationship.
Why
Google News forbids concealed or misrepresented sponsored content.

Is material automation or generative-AI use disclosed when readers would reasonably expect it?

Priority: 4
What
Evidence-backed tactic: Explain where automation contributed and what human checks occurred.
How
add a creation-method note covering generation, data, review, and limitations.
Why
Google recommends context about how automatically generated content was created.

Are conflicts of interest and relevant affiliations disclosed at claim level?

Priority: 4
What
Evidence-backed tactic: Reveal relationships that could affect interpretation.
How
collect author/reviewer disclosures and place material conflicts near the content.
Why
Transparent incentives help readers evaluate recommendations.

Does high-stakes advice use qualified review and authoritative local sources?

Priority: 5
What
Evidence-backed tactic: Apply stronger evidence and review to health, legal, financial, safety, and civic guidance.
How
require jurisdiction-appropriate primary sources, specialist review, update cadence, and escalation language.
Why
Error cost is high and rules change.

Is there a governed registry for priority people, organizations, products, places, and concepts?

Priority: 5
What
Evidence-backed tactic: Maintain one internal record per entity.
How
assign a stable ID, canonical name, type, URL, aliases, relationships, source, owner, and review date.
Why
A registry prevents contradictory identities across pages and feeds.

Is each entity's canonical public name consistent across the site?

Priority: 5
What
Evidence-backed tactic: Use the same primary identity in visible text, metadata, structured data, profiles, and feeds.
How
compare templates and data sources against the entity registry.
Why
Google recommends Organization names consistent with the site name.

Are legitimate aliases, handles, abbreviations, and former names recorded?

Priority: 4
What
Evidence-backed tactic: Help readers reconcile names without treating every spelling as a separate entity.
How
show relevant aliases in profiles and use `alternateName` where supported.
Why
Google exposes `alternateName` for profile and organization identity.

Are ambiguous names explicitly disambiguated?

Priority: 5
What
Evidence-backed tactic: Distinguish entities sharing a name or acronym.
How
add type, location, affiliation, role, version, and canonical link at first mention.
Why
Names alone may not uniquely identify a person, company, product, or concept.

Does each important entity have one stable canonical URL?

Priority: 5
What
Evidence-backed tactic: Give the entity a durable web identity.
How
choose the authoritative page, keep it current, link to it consistently, and redirect retired URLs.
Why
Google says an Organization URL helps uniquely identify it.

Does the organization's canonical page state accurate identity and operating details?

Priority: 5
What
Evidence-backed tactic: Cover name, description, URL, logo, contact, address, legal identity, and relevant external profiles.
How
maintain these facts on the home or About page and reconcile them with markup.
Why
Google uses Organization data to understand administrative details and disambiguation.

Does each marked-up profile focus on one affiliated Person or Organization?

Priority: 4
What
Confirmed platform requirement: Keep the profile's primary entity unambiguous.
How
use `ProfilePage.mainEntity` for that one person or organization and show the same focus visibly.
Why
Google lists single-entity focus as a ProfilePage content guideline.

Do Article author objects link to their canonical profile URLs?

Priority: 4
What
Evidence-backed tactic: Connect each work to a stable creator identity.
How
set `author.url` to the internal author page and use ProfilePage markup there when appropriate.
Why
Google recommends author URLs to uniquely identify creators.

Are `sameAs` links limited to high-confidence identity matches?

Priority: 4
What
Evidence-backed tactic: Link only external pages that unambiguously represent the same entity.
How
verify ownership, name, URL, and current status. Remove stale or fan-created profiles.
Why
Because `sameAs` is an identity assertion, keep generic related links elsewhere.

Are official identifiers included only when verified and correctly scoped?

Priority: 3
What
Evidence-backed tactic: Use internal IDs, LEI, VAT, tax, GS1, ROR, DOI, or other identifiers for the right entity type.
How
validate against the issuing registry and record jurisdiction and format.
Why
Identifiers can disambiguate entities across systems.

Are parent, subsidiary, brand, product, founder, employee, and ownership relationships explicit?

Priority: 4
What
Evidence-backed tactic: Model how entities relate rather than relying on name proximity.
How
state relationships visibly and encode supported properties consistently.
Why
Clear relationships prevent a model or reader from conflating a brand, legal entity, and product.

Are people, organizations, products, events, places, and concepts typed correctly?

Priority: 5
What
Evidence-backed tactic: Avoid one generic entity object for unlike things.
How
audit visible nouns and JSON-LD types against the entity registry and source evidence.
Why
Correct typing preserves meaning and property validity.

Does markup distinguish what a work is about from entities it merely mentions?

Priority: 4
What
Evidence-backed tactic: Separate the primary subject from incidental references.
How
use `about` or `mainEntity` for subject matter and `mentions` only for secondary referenced entities.
Why
Schema.org defines these relationships differently.

Are expertise-topic assertions narrow, verifiable, and non-promotional?

Priority: 2
What
Experimental practice: Describe what a Person or Organization knows about without claiming unsupported mastery.
How
use visible evidence and, if useful, `knowsAbout` links or terms tied to real work.
Why
Schema.org says the property suggests possible expertise but does not imply it.

Are entity facts consistent across text, schema, images, video, and downloadable data?

Priority: 5
What
Evidence-backed tactic: Prevent cross-format contradictions in names, specs, dates, pricing, and claims.
How
compare all renditions against the canonical entity record during release.
Why
Microsoft recommends alignment across formats to reduce ambiguity.

Are local business name, address, phone, hours, and service facts current everywhere?

Priority: 5
What
Evidence-backed tactic: Maintain one source of truth for local operating details.
How
reconcile website, markup, Bing Places, Google Business Profile, directories, and location pages.
Why
Microsoft highlights current address, hours, and contact information for location-based AI answers.

Are current, legal, and former organization names distinguished while mergers, acquisitions, and rebrands are documented with dates?

Priority: 2
What
Experimental practice: Preserve identity continuity without conflating a current organization with former entities or names.
How
Keep `name`, `legalName`, `alternateName`, identifiers, and `sameAs` links consistent. Document dated changes visibly and update redirects and relationships.
Why
Consistent identity fields aid organization disambiguation. Treat dated corporate-history modeling as an operational inference because it is not a documented generative-search ranking factor.

Are product editions, models, plans, and versions unambiguously separated?

Priority: 5
What
Evidence-backed tactic: Prevent specifications or prices from crossing versions.
How
give each material variant a stable identifier, lifecycle dates, compatible versions, and canonical detail page.
Why
Entity ambiguity can produce wrong recommendations.

Do internal links name the destination entity or topic descriptively?

Priority: 4
What
Evidence-backed tactic: Replace generic anchors with concise, contextual identity labels.
How
make important entity pages reachable through real anchors whose text makes sense out of context.
Why
Google recommends descriptive, relevant anchor text.

Are external identity sources periodically reconciled against first-party facts?

Priority: 4
What
Evidence-backed tactic: Find conflicts in trusted registries, profiles, reviews, and knowledge panels.
How
monitor high-impact external records, document the authoritative source, and correct owned surfaces or request fixes.
Why
AI answers can synthesize information from the brand site alongside other sources.

Are only documented applicable schema types used, without inventing a “GEO schema”?

Priority: 5
What
Confirmed platform requirement: Choose the most specific supported vocabulary for the real page.
How
validate against platform documentation and schema.org.
Why
Structured data supplies explicit clues but no special answer-engine type is documented.

Is JSON-LD syntactically valid, reachable, and attached to the page it describes?

Priority: 4
What
Confirmed platform requirement: Use a maintainable supported format.
How
parse production markup and test required properties and URLs.
Why
Google recommends JSON-LD and requires access to the marked page.

Are required and useful recommended properties complete, accurate, and current?

Priority: 5
What
Confirmed platform requirement: Avoid partial or stale machine facts.
How
validate against the canonical database and expiry rules on every release.
Why
Incomplete required properties lose feature eligibility. Stale time-sensitive data may not display.

Are image and media URLs referenced by structured data crawlable and indexable?

Priority: 4
What
Confirmed platform requirement: Ensure machines can retrieve declared assets.
How
test status, robots, canonical, MIME, and stable URL from the edge.
Why
Google cannot use inaccessible referenced images.

Is paywalled content declared with accurate platform-supported markup rather than crawler-only cloaking?

Priority: 5
What
Confirmed platform requirement: Distinguish restricted access honestly.
How
implement Google `isAccessibleForFree`/`hasPart` where applicable and Apple's page-level signal.
Why
Correct markup preserves eligibility while communicating restrictions.

Are visual relationships also programmatically determinable?

Priority: 5
What
Evidence-backed tactic: Encode headings, lists, definitions, tables, landmarks, and labels with appropriate HTML.
How
inspect DOM and accessibility-tree output, including content injected by components.
Why
Semantic relationships remain available when presentation changes.

Are citations and internal references real crawlable links?

Priority: 5
What
Evidence-backed tactic: Use `<a href>` rather than click handlers, styled spans, or script-only navigation.
How
inspect rendered HTML and test links without client JavaScript.
Why
Links expose a durable relationship and destination.

Are self-contained charts, diagrams, code listings, and quotations grouped with `figure` and `figcaption`?

Priority: 3
What
Evidence-backed tactic: Bind an asset to its caption and stable label.
How
use semantic figure markup and refer to it by label rather than “above” or “below.”
Why
WHATWG defines figures as self-contained units with optional captions.

Are block and inline quotations marked and attributed correctly?

Priority: 3
What
Evidence-backed tactic: Distinguish quoted material from author prose.
How
use `blockquote` or `q` where appropriate, a valid source URL when useful, and visible attribution outside the quote.
Why
HTML semantics preserve quotation boundaries.

Is the `cite` element used for a work title rather than a person or quotation?

Priority: 3
What
Evidence-backed tactic: Keep citation semantics valid.
How
mark only the title of the referenced work with `cite` and provide a separate clickable source.
Why
WHATWG explicitly limits `cite` to titles of works.

Are defining instances encoded and linkable where definitions matter?

Priority: 3
What
Evidence-backed tactic: Make a term and its definition explicitly related.
How
use `dfn`, a stable fragment ID, and contextual links from later uses when helpful.
Why
WHATWG defines `dfn` as the defining instance of a term.

Are publication, modification, event, and expiry dates machine-readable and correctly labeled?

Priority: 4
What
Evidence-backed tactic: Separate different temporal facts.
How
use visible labels, valid `<time datetime>` values, correct timezone where relevant, and matching structured data.
Why
The time element provides a machine-readable representation.

Does structured data describe content users can actually see?

Priority: 5
What
Confirmed platform requirement: Prevent hidden, misleading, or contradictory claims in JSON-LD.
How
compare every material property with the canonical rendered page and source data.
Why
Google requires markup to represent visible page content.

Is the most specific accurate, Google-supported page type used for the intended feature?

Priority: 4
What
Evidence-backed tactic: Avoid generic or invented schema where a documented type applies.
How
start from the relevant Google feature guide, then add broader Schema.org only for truthful secondary meaning.
Why
Google says Search documentation is definitive for Google behavior.

Are fewer complete, accurate properties favored over broad incomplete markup?

Priority: 4
What
Evidence-backed tactic: Prioritize data quality.
How
populate required and meaningful recommended properties from governed sources. Omit unknown values.
Why
Google recommends supplying fewer complete and accurate properties rather than every possible property.

Do Article author, date, headline, and image values match the article?

Priority: 4
What
Confirmed platform requirement: Keep creator and publication metadata aligned.
How
validate every author separately, stable author URLs, accurate dates, and representative crawlable images.
Why
Google's Article guide defines the supported fields and author best practices.

Is Organization markup placed on the canonical home or About page rather than repeated indiscriminately?

Priority: 3
What
Evidence-backed tactic: Centralize identity details.
How
maintain one complete Organization node and reference its stable `@id` from other markup.
Why
Google recommends organization details on a single home or About page.

Are first-party datasets described with discoverable metadata?

Priority: 3
What
Evidence-backed tactic: Publish name, description, creator, identifier, license, coverage, distribution, and provenance.
How
add Dataset/DataDownload markup to the canonical dataset landing page and validate it.
Why
Google uses structured metadata for Dataset Search discovery.

Are `citation` and `isBasedOn` used only when they truthfully express work lineage?

Priority: 1
What
Experimental practice: Encode a reference or derivation in addition to visible citations.
How
add these Schema.org properties to CreativeWork objects and validate consumers before scaling.
Why
The vocabulary can make provenance relationships explicit.

Is structured data validated without being sold as an AI citation switch?

Priority: 5
What
Evidence-backed tactic: Test syntax, feature eligibility, visible consistency, and post-deploy errors.
How
use validator and platform reports, then measure outcomes.
Why
Google says structured data is not required for generative AI search and there is no special AI schema.

Do images and videos add evidence or explanation rather than decoration alone?

Priority: 4
What
Evidence-backed tactic: Use media that helps answer the user's need.
How
map each asset to a claim, procedure, comparison, or demonstration and remove redundant stock imagery.
Why
Google recommends high-quality relevant media for generative AI search opportunities.

Does every informative image have useful contextual alt text?

Priority: 5
What
Evidence-backed tactic: Convey the image's purpose and the information needed to understand it.
How
write concise alt text in page context. Use empty alt for decorative images and avoid keyword stuffing.
Why
Google calls alt text its most important image metadata and WCAG requires text alternatives.

Is each important image placed near explanatory text and a factual caption?

Priority: 4
What
Evidence-backed tactic: Connect media to the surrounding claim.
How
place it in the relevant section, name entities, date the asset when material, and cite its source.
Why
Google derives image subject matter from nearby page content, captions, and titles.

Do charts and diagrams have a complete text or data equivalent?

Priority: 5
What
Evidence-backed tactic: Expose trends, labels, units, and conclusions outside pixels.
How
provide concise alt plus nearby interpretation, a data table, or long description when alt is insufficient.
Why
WCAG says complex non-text content needs an alternative that presents equivalent information.

Are figures given stable captions and referenced by label rather than position?

Priority: 3
What
Evidence-backed tactic: Preserve meaning across responsive layouts and extraction.
How
use `figure`, `figcaption`, an ID, and references such as “Figure 2.”
Why
WHATWG recommends labels instead of “above” or “below” references.

Is the preferred preview image representative, specific, and high quality?

Priority: 4
What
Evidence-backed tactic: Avoid generic logos or misleading thumbnails.
How
select a relevant image, expose consistent image metadata, and review crops and aspect ratios.
Why
Google recommends representative, non-generic, high-resolution preferred images.

Are image filenames, alt text, captions, and embedded text localized?

Priority: 3
What
Evidence-backed tactic: Adapt the whole asset for its audience.
How
translate meaningful metadata and recreate text-heavy graphics rather than overlaying partial translations.
Why
Google recommends translating localized image filenames and contextual metadata.

Do important media assets use stable, crawlable canonical URLs?

Priority: 4
What
Evidence-backed tactic: Preserve discoverability and attribution across reuse.
How
avoid expiring signed URLs for public assets, redirect replacements, and use the same image URL when the same asset recurs.
Why
Google recommends consistent image URLs and stable video URLs.

Are creator, credit, copyright, license, and acquisition details attached to reusable assets?

Priority: 4
What
Evidence-backed tactic: Preserve attribution and rights.
How
add IPTC or structured metadata, a visible credit, license URL, and licensor route where relevant.
Why
Google can show creator and licensing details in Images.

Are AI-generated ecommerce images and product data labeled with the required metadata?

Priority: 5
What
Confirmed platform requirement: Identify synthetic product media and generated catalog fields.
How
embed IPTC `DigitalSourceType` `TrainedAlgorithmicMedia` for AI images and submit generated product fields separately as labeled AI content.
Why
Google Merchant Center policy requires this treatment.

Where the workflow supports it, do high-risk original media assets carry verifiable Content Credentials?

Priority: 2
What
Evidence-backed tactic: Add cryptographically verifiable, tamper-evident provenance for origin and edits.
How
Create a C2PA 2.4 manifest, bind and sign it, preserve validated ingredient provenance through authorized transformations, and validate the delivered asset.
Why
C2PA standardizes signed provenance assertions, content bindings, and validation states for media workflows.

Does each primary video have a dedicated, stable watch page when appropriate?

Priority: 3
What
Evidence-backed tactic: Give a video its own canonical context.
How
create a page where that video is the main content with a unique title and description.
Why
Google requires a dedicated watch page for eligibility in video features.

Do video metadata and markup accurately describe the actual video?

Priority: 4
What
Confirmed platform requirement: Align title, description, thumbnail, dates, content URL, and regions with the media.
How
validate VideoObject and compare every field with the playable asset.
Why
Google requires structured data to be consistent with video content and other metadata.

Do prerecorded audio and video provide accurate captions and descriptive transcripts?

Priority: 5
What
Evidence-backed tactic: Expose speech, relevant sounds, and visual information needed for understanding in text.
How
human-review captions, publish a transcript, identify speakers, and include descriptions needed to understand visuals.
Why
W3C identifies captions and transcripts as accessibility alternatives.

Are video chapters, timestamps, transcript claims, and surrounding page text aligned?

Priority: 4
What
Evidence-backed tactic: Make key moments navigable without cross-format contradiction.
How
add accurate chapter labels or Clip/SeekToAction data and reconcile dates, names, numbers, and conclusions across formats.
Why
Google supports key moments and Microsoft recommends cross-format entity consistency.

Are locale alternates correctly annotated and mutually coherent?

Priority: 5
What
Confirmed platform requirement: Connect language/region equivalents.
How
validate `hreflang` codes, self/return references, canonicals, and sitemap or HTML annotations.
Why
It helps Google select the right locale URL.

Are robots, meta, status, canonical, and structured-data rules consistent across locale variants?

Priority: 5
What
Confirmed platform requirement: Prevent one language from silently losing eligibility.
How
run the full retrieval matrix for every locale.
Why
Google explicitly asks locale-adaptive sites to apply robots controls consistently.

Is important localized content available without IP or `Accept-Language` dependence and without forced redirects?

Priority: 5
What
Confirmed platform requirement: Serve stable locale URLs to any valid crawler.
How
test from US and target regions with missing/different language headers.
Why
Googlebot often uses US IPs and sends no `Accept-Language`.

Does each language version have a distinct stable URL?

Priority: 5
What
Evidence-backed tactic: Make localized content independently addressable.
How
use locale-specific paths or hosts and avoid cookie-only or header-only language swaps.
Why
Google recommends different URLs because dynamic variations may not all be crawled.

Do all `hreflang` variants list themselves and every reciprocal version with absolute URLs?

Priority: 5
What
Confirmed platform requirement: Keep the locale cluster complete and bidirectional.
How
generate one consistent set, validate language-region codes, and test return links.
Why
Google may ignore non-reciprocal or incomplete annotations.

Is an `x-default` fallback declared for unmatched language or region users?

Priority: 4
What
Evidence-backed tactic: Identify the neutral selector or default page.
How
add `hreflang="x-default"` to the same alternate set and verify its purpose.
Why
Google documents `x-default` for users whose locale is not explicitly targeted.

Does each page use one clear primary visible language for content and navigation?

Priority: 5
What
Evidence-backed tactic: Avoid mixed-language pages and boilerplate-only translation.
How
audit main content, menus, widgets, errors, captions, and structured text.
Why
Google determines language from visible content and recommends a single language per page.

Is localized content adapted by a fluent human for local intent and conventions?

Priority: 5
What
Evidence-backed tactic: Preserve meaning rather than mirror sentence structure mechanically.
How
review terminology, examples, tone, legal context, sources, and user tasks with a locale expert.
Why
A translated shell around unchanged content is a poor user experience.

Can users switch language or region without forced automatic redirection?

Priority: 5
What
Evidence-backed tactic: Preserve user control and linkability.
How
provide visible alternate links, remember a preference without hiding URLs, and avoid IP- or language-forced redirects.
Why
Google warns automatic redirects can prevent users and crawlers from seeing variants.

Is the document language declared with a valid `lang` value on the root element?

Priority: 5
What
Confirmed platform requirement: Expose the page's default human language.
How
use a valid BCP 47 tag on `<html>` and test generated layouts.
Why
W3C identifies the `lang` attribute as the standard declaration for text processing and accessibility.

Are inline language changes and bidirectional text marked correctly?

Priority: 4
What
Confirmed platform requirement: Preserve pronunciation, shaping, and reading order.
How
add `lang` to passages in another language and use appropriate `dir`, `bdi`, or `bdo` handling.
Why
W3C guidance requires declaration at the highest applicable level and direction separately.

Are currency, tax, units, dates, time zones, availability, and examples locally correct?

Priority: 5
What
Evidence-backed tactic: Adapt facts that change the answer by market.
How
source values per locale, display units and conversion basis, and date volatile details.
Why
Location is part of answer relevance and ChatGPT can use approximate or precise location.

Do legal, regulatory, medical, and civic claims cite authoritative sources for the target jurisdiction?

Priority: 5
What
Evidence-backed tactic: Prevent one country's rule from being generalized globally.
How
name jurisdiction, effective date, issuing authority, and exceptions near the claim.
Why
Local context materially changes high-stakes answers.

Are local address, phone, hours, service area, and contact methods consistent?

Priority: 5
What
Evidence-backed tactic: Keep location entities accurate in every locale.
How
reconcile page text, schema, platform listings, maps, and local profiles from one governed source.
Why
Google lists local addresses, phone numbers, currency, and local links among locale signals.

Are local names, transliterations, grammatical forms, and aliases reconciled to the same entity?

Priority: 4
What
Evidence-backed tactic: Connect language-specific labels without erasing native usage.
How
maintain locale-aware canonical names and alternate names, and link to the same stable entity ID.
Why
Entity resolution can fail when translations look like separate entities.

Does each localized page cite sources credible and accessible to that audience?

Priority: 4
What
Evidence-backed tactic: Avoid importing all evidence from another language or market.
How
prefer authoritative local primary sources, retain global sources where applicable, and explain cross-market transfer.
Why
Search and AI source ecosystems vary by language and geography.

Are priority prompts tested across language, country, city, and personalization contexts?

Priority: 2
What
Experimental practice: Detect where answers, citations, and entity matching diverge.
How
run named locales, natural local-language paraphrases, repeated dates, and controlled locations. Record null results.
Why
OpenAI documents location-aware rewriting and research finds cross-language variation.

For same-language regional duplicates, are canonical and `hreflang` signals coordinated?

Priority: 5
What
Confirmed platform requirement: Choose a preferred duplicate while preserving regional targeting.
How
point each regional version to the intended canonical and maintain the complete alternate set.
Why
Google explicitly recommends canonical plus `hreflang` for same-language regional duplicates.

Is each target page indexed and eligible to appear with a snippet in Google Search?

Priority: 5
What
Confirmed platform requirement: Establish baseline AI-feature eligibility.
How
Inspect canonical/index status and snippet controls in Search Console.
Why
This is Google's stated technical gate for supporting links.

Can Googlebot crawl every target page and required resource?

Priority: 5
What
Confirmed platform requirement: Allow Search crawling where AI visibility is desired.
How
Test robots, HTTP response, rendered HTML, CSS, JS, and media.
Why
Googlebot is the access control for AI features inside Search.

Is `nosnippet` absent where Google AI citation eligibility is desired?

Priority: 5
What
Confirmed platform requirement: Audit page and header directives.
How
Check HTML and `X-Robots-Tag` on canonical and duplicate responses.
Why
`nosnippet` prevents direct input to AI Overviews and AI Mode.

Is `max-snippet` intentionally sized for the desired AI-preview policy?

Priority: 4
What
Confirmed platform requirement: Avoid accidental zero or overly restrictive values.
How
Inventory effective directives after proxy/CDN composition.
Why
Google applies the limit to direct input for AI Overviews and AI Mode.

Is `data-nosnippet` applied only to content intentionally excluded from Google AI previews?

Priority: 4
What
Confirmed platform requirement: Protect selected text without suppressing the entire page.
How
Put the attribute on valid `div`, `span`, or `section` elements in initial/rendered DOM.
Why
It controls text-level preview use.

Is `noindex` absent from pages intended for Google AI features?

Priority: 5
What
Confirmed platform requirement: Check HTML and non-HTML headers.
How
Compare origin, CDN, mobile, and rendered responses.
Why
A page must remain eligible for Search indexing.

Is Google-Extended treated separately from Google Search AI features?

Priority: 5
What
Confirmed platform requirement: Do not use it as an AI Overviews/AI Mode switch.
How
Set it only for Gemini Apps/Vertex grounding and future Gemini training policy.
Why
It does not govern Search inclusion.

Is GoogleOther excluded from Google Search visibility decisions?

Priority: 3
What
Confirmed platform requirement: Avoid calling it the AI Overviews crawler.
How
Classify it as generic product-team/R&D fetching unless a specific official use is documented.
Why
Its rules affect no specific product.

Are Google inspection-tool results distinguished from production Googlebot access?

Priority: 4
What
Confirmed platform requirement: Test both tooling and actual crawl/index evidence.
How
Avoid WAF rules that only permit `Google-InspectionTool`.
Why
That token affects tests but does not affect Search.

Is suspected Googlebot traffic verified beyond its user-agent string?

Priority: 5
What
Confirmed platform requirement: Defend against spoofing without blocking genuine crawls.
How
Use reverse DNS or Google's published ranges.
Why
User-agent strings are spoofable.

Does the mobile response contain equivalent primary content, metadata, links, and structured data?

Priority: 5
What
Confirmed platform requirement: Audit what smartphone Googlebot sees.
How
Compare rendered mobile and desktop DOMs and response directives.
Why
Google primarily indexes mobile content.

Are related subtopics independently crawlable for query fan-out?

Priority: 4
What
Evidence-backed tactic: Make supporting pages and sections discoverable through normal links.
How
Crawl the topic cluster from its hub and inspect canonicals/index state.
Why
AI Mode and AI Overviews may issue multiple related searches.

Is Google generative-AI visibility measured in the correct Search Console views?

Priority: 4
What
Confirmed platform requirement: Use the dedicated report when available and retain the overall Web report.
How
Record rollout availability and export comparable date ranges.
Why
AI feature data remains included in overall performance.

Are URL, country, device, and time dimensions reviewed in Google's generative-AI report?

Priority: 3
What
Confirmed platform requirement: Separate visibility shifts by surface context.
How
Monitor pages, countries, devices, and hourly/daily/weekly/monthly trends available to the property.
Why
Aggregate totals can hide regional or template failures.

After preview-control changes, was recrawl and processing time allowed and verified?

Priority: 5
What
Confirmed platform requirement: Do not call a change ineffective immediately.
How
inspect fetched HTML, request recrawl, and monitor until the canonical is reprocessed.
Why
Google says changes can take days to months to recrawl.

Has the team rejected any alleged special Google “GEO requirement”?

Priority: 5
What
Confirmed platform requirement: Keep Google AI eligibility grounded in Search technical requirements and policies.
How
Require an official source before adding a special tag, file, or markup.
Why
Google states there are no additional technical requirements or special optimizations.

Is OAI-SearchBot allowed on every page intended for ChatGPT Search summaries and citations?

Priority: 5
What
Confirmed platform requirement: Permit the search-specific bot.
How
Test the effective robots group for canonical content and assets.
Why
Opted-out sites are not shown in Search answers, though navigational links may remain.

Does the CDN/WAF allow current OAI-SearchBot IP ranges as well as its user agent?

Priority: 5
What
Confirmed platform requirement: Remove network-layer false blocks.
How
consume `https://openai.com/searchbot.json`, combine IP and UA conditions, and test origin logs.
Why
OpenAI explicitly requires host/CDN access from published ranges.

Do user-agent rules tolerate OAI-SearchBot version changes and the robots marker?

Priority: 4
What
Confirmed platform requirement: Match the stable product token instead of a frozen full string.
How
test normal and `robots.txt`-marker request forms.
Why
OpenAI says version numbers may change and robots fetches may carry an extra marker.

Is GPTBot policy independent from OAI-SearchBot policy?

Priority: 5
What
Confirmed platform requirement: Decide training use without accidentally suppressing Search.
How
give each token its own robots group and tests.
Why
OpenAI permits Search while disallowing model-training crawling.

Is ChatGPT-User excluded from automatic Search inclusion logic?

Priority: 4
What
Confirmed platform requirement: Do not use it as the Search allow/deny token.
How
manage Search through OAI-SearchBot and treat ChatGPT-User as user-triggered access.
Why
OpenAI says ChatGPT-User does not determine Search appearance.

Has the site tested user-triggered ChatGPT access separately from robots-controlled crawling?

Priority: 4
What
Confirmed platform requirement: Observe actions from ChatGPT-User.
How
test representative public URLs and inspect authenticated/WAF behavior.
Why
robots.txt may not apply to user-initiated actions.

Was the documented robots-policy propagation window considered after an OpenAI change?

Priority: 4
What
Confirmed platform requirement: Avoid premature pass/fail judgments.
How
retest after at least the stated adjustment window and confirm live requests.
Why
OpenAI says Search systems may take about 24 hours to adjust.

If even title/link exposure is prohibited, is `noindex` used on a crawlable page rather than relying only on OAI-SearchBot disallow?

Priority: 5
What
Confirmed platform requirement: Align the control with the removal goal.
How
permit the bot to read `noindex`, then verify de-indexing.
Why
OpenAI may surface a disallowed URL/title learned elsewhere.

Is ChatGPT referral traffic identified by its documented campaign parameter?

Priority: 4
What
Confirmed platform requirement: Preserve and report `utm_source=chatgpt.com`.
How
prevent redirects from stripping it and segment landing/conversion analytics.
Why
OpenAI adds this parameter to referral URLs.

Are intended ChatGPT Search pages publicly reachable without login, CAPTCHA, or consent dead ends?

Priority: 5
What
Confirmed platform requirement: Ensure a usable public response.
How
fetch through CDN/origin with the verified bot path and a clean session.
Why
OpenAI describes public websites as eligible and separately flags auth/challenge blocks.

Is ChatGPT Search tested against exact branded prompts as well as likely rewrites and follow-up queries?

Priority: 4
What
Evidence-backed tactic: Validate discovery through query variants.
How
maintain a prompt set covering subtopics, recency, comparison, and locale.
Why
ChatGPT may rewrite one prompt into multiple targeted partner queries.

For commerce sites, is current first-party product data supplied through supported integrations or feeds where eligible?

Priority: 4
What
Evidence-backed tactic: Reduce stale product facts.
How
validate Shopify Catalog integration or apply for direct product-feed access and reconcile feed-to-page data.
Why
OpenAI says direct feeds help reflect current products.

Do commerce fields such as availability, price, descriptions, and merchant identity agree across feed, structured metadata, and landing page?

Priority: 4
What
Confirmed platform requirement: Eliminate conflicting product facts.
How
diff catalog exports against canonical pages on every update.
Why
ChatGPT uses merchant and third-party metadata and ranks merchant options partly on availability and price.

Are interactive pages semantically operable for ChatGPT Agent in Atlas?

Priority: 3
What
Evidence-backed tactic: Make buttons, menus, forms, roles, labels, and states machine-readable.
How
audit against WAI-ARIA patterns and run representative agent journeys.
Why
OpenAI says Atlas uses ARIA semantics to interpret interfaces.

Can Bingbot crawl target content and resources?

Priority: 5
What
Confirmed platform requirement: Establish Bing index eligibility feeding Bing and AI-powered experiences.
How
test robots, live fetch, response, and rendered markup in Bing Webmaster Tools.
Why
Bing ties Copilot discovery to crawl/index availability.

If a specific Bingbot group exists, does it repeat needed generic directives?

Priority: 4
What
Confirmed platform requirement: Avoid unintended policy gaps.
How
parse the effective Bingbot group rather than assuming `*` merges into it.
Why
Bing says a specific group causes it to ignore generic directives.

Is Bingbot allowed to crawl pages carrying `noindex` until Bing can observe the directive?

Priority: 5
What
Confirmed platform requirement: Use the correct removal mechanism.
How
remove robots disallow, serve `noindex`, and monitor index state.
Why
Bing must fetch the page to read `noindex`.

Is Microsoft's generative-model training preference deliberately configured rather than inferred from crawl access?

Priority: 4
What
Confirmed platform requirement: Audit Bing-supported meta/X-Robots controls, including the currently documented training-use directive.
How
record the exact effective tag and verify in fetched HTML/header.
Why
Bing documents content-display and generative-AI data-use controls separately.

If IndexNow-participating engines matter in target markets, are new, meaningfully changed, and deleted URLs submitted?

Priority: 3
What
Evidence-backed tactic: Use IndexNow as an optional freshness notification alongside crawlable links and XML sitemaps.
How
Enable a trusted CMS or CDN integration, or deploy a verified API key, and submit only recent lifecycle changes.
Why
Participating engines can receive change notifications sooner, but each engine still decides whether and when to crawl or index a URL.

Are IndexNow and XML sitemaps used together rather than treated as substitutes?

Priority: 5
What
Evidence-backed tactic: Pair real-time change notification with comprehensive inventory.
How
compare submitted-change logs with sitemap coverage.
Why
Bing explicitly recommends the combined signals for AI-powered search.

Are target URLs checked with both indexed and Live URL views in Bing URL Inspection?

Priority: 5
What
Confirmed platform requirement: Separate stored index state from current fetch state.
How
inspect discovery, crawl, HTTP, HTML, canonical, markup, and live response.
Why
The tool shows what Bingbot sees and why a URL is excluded.

Is Bing crawl throttling configured through supported controls rather than error responses?

Priority: 4
What
Confirmed platform requirement: Protect capacity without destroying access.
How
use Crawl Control or a deliberate `crawl-delay`, then monitor freshness.
Why
Bing honors crawl-delay ahead of dashboard settings.

Is PerplexityBot allowed where Perplexity search visibility is desired?

Priority: 5
What
Confirmed platform requirement: Permit the automatic search-index crawler.
How
test its robots group across canonical content and resources.
Why
Perplexity recommends allowing it to appear in results.

Does the WAF allow verified PerplexityBot traffic using both user-agent and IP?

Priority: 5
What
Confirmed platform requirement: Prevent security controls from nullifying robots intent.
How
combine UA matching with the official `perplexitybot.json` ranges.
Why
Perplexity explicitly recommends the combined condition.

Are Perplexity IP sets refreshed automatically from official endpoints?

Priority: 4
What
Confirmed platform requirement: Keep Cloudflare/AWS rules current.
How
periodically fetch both bot and user JSON sources, validate, stage, and alert on deltas.
Why
Perplexity says the addresses update regularly.

Is Perplexity-User treated separately from PerplexityBot?

Priority: 4
What
Confirmed platform requirement: Audit user-requested page fetches independently.
How
test both identities and document desired access.
Why
Perplexity-User supports question-time actions and is not the automatic index crawler.

Was Perplexity's stated policy-propagation delay allowed after crawler-rule changes?

Priority: 4
What
Confirmed platform requirement: Avoid testing too early.
How
repeat validation after the documented window and confirm logs.
Why
Perplexity says changes may take up to 24 hours.

Is PerplexityBot absent from the model-training opt-out matrix?

Priority: 5
What
Confirmed platform requirement: Do not mislabel its purpose.
How
classify it as search indexing and handle training policy elsewhere.
Why
Perplexity states the bot is not used to crawl for foundation models.

If all Perplexity mention is prohibited, is robots disallow supplemented by an appropriate removal control?

Priority: 4
What
Confirmed platform requirement: Account for residual domain/headline/summary indexing.
How
verify current displayed residue and escalate through publisher support where needed.
Why
Perplexity says blocked sites may still have limited facts indexed.

Is Claude-SearchBot allowed where Claude search visibility is desired?

Priority: 5
What
Confirmed platform requirement: Permit the search-quality crawler.
How
test the exact token across every host and key path.
Why
Anthropic says disabling it may reduce visibility and accuracy in user search results.

Is Claude-User access deliberately configured and tested?

Priority: 4
What
Confirmed platform requirement: Decide whether Claude may retrieve pages at a user's direction.
How
test representative public, paywalled, and protected paths.
Why
Disabling it prevents user-initiated retrieval and may reduce visibility.

Is ClaudeBot training policy independent from Claude-SearchBot and Claude-User?

Priority: 5
What
Confirmed platform requirement: Separate model-development collection from answer retrieval.
How
maintain three explicit robots groups.
Why
Anthropic documents a different purpose and effect for each.

Are Anthropic rules present on every opted-in or opted-out subdomain?

Priority: 5
What
Confirmed platform requirement: Avoid partial policy coverage.
How
fetch `/robots.txt` and target responses host by host.
Why
Anthropic explicitly asks publishers to configure every subdomain.

If Anthropic crawl load needs control, is its supported non-standard `crawl-delay` used carefully?

Priority: 3
What
Confirmed platform requirement: Throttle rather than block or error.
How
set a modest token-specific value and inspect freshness and logs.
Why
Anthropic says its bots support crawl-delay where appropriate.

Are Anthropic bot source IPs verified against the current official range list?

Priority: 5
What
Confirmed platform requirement: Use Anthropic's published bot ranges as an authentication signal rather than assuming a static list.
How
fetch the official JSON on a controlled schedule, match the declared agent and source range, and alert on changes.
Why
Anthropic now publishes a living bot range list and warns that IP blocking is not a persistent opt-out method.

Is `noindex` used when content must not appear in Claude web-search outputs?

Priority: 5
What
Confirmed platform requirement: Signal search partners not to index the page.
How
serve a crawlable `noindex` and verify disappearance over time.
Why
Anthropic documents this as an all-content-type exclusion control.

Is confidential content protected by authentication rather than crawler etiquette?

Priority: 5
What
Confirmed platform requirement: Enforce access at the application boundary.
How
require authorization and remove leaked public URLs. Use the owner-verified removal route if already surfaced.
Why
Anthropic recommends password protection for private material.

Is Meta-WebIndexer explicitly allowed to crawl every public URL class intended for Meta AI discovery?

Priority: 5
What
Evidence-backed tactic: Maintain a token-specific robots policy for public canonical content.
How
Publish a `User-agent: meta-webindexer` group, allow intended public routes, exclude nonpublic routes, and compare the production robots response with request logs.
Why
Meta says this crawler supports Meta AI result relevance and accuracy and that allowing it helps Meta cite and link to content.

Is Meta-ExternalAgent governed separately from Meta-WebIndexer under the training and indexing policy?

Priority: 5
What
Evidence-backed tactic: Make an explicit allow-or-disallow decision for Meta-ExternalAgent.
How
Publish a separate `User-agent: meta-externalagent` group, document its owner and permitted URL scope, and review it independently from WebIndexer.
Why
Meta assigns ExternalAgent a different purpose: foundation-model training or product improvement through direct indexing.

Do private and sensitive routes remain protected if Meta-ExternalFetcher bypasses robots.txt?

Priority: 5
What
Evidence-backed tactic: Treat robots.txt as a crawl preference for user-requested fetches. It is not an authorization boundary.
How
Require authentication and server-side authorization for every nonpublic resource, remove secrets from public URLs and markup, and test protected routes while logged out.
Why
Meta says ExternalFetcher retrieves individual links at a user's request and may bypass robots.txt.

Are Meta robots.txt changes evaluated after the documented cache window?

Priority: 4
What
Evidence-backed tactic: Account for a propagation period of up to 24 hours.
How
Record publication time, avoid contradictory edits during the window, compare crawler requests before and after it, and do not declare a rule ineffective prematurely.
Why
Meta says its crawlers may cache robots.txt for up to 24 hours.

Can logs distinguish Meta-WebIndexer, Meta-ExternalAgent, and Meta-ExternalFetcher traffic?

Priority: 4
What
Evidence-backed tactic: Preserve a normalized crawler identity beside the raw User-Agent.
How
Classify the documented tokens separately and retain time, URL, status, bytes, latency, source IP, and verification state.
Why
Meta documents different purposes and User-Agent strings for these crawlers.

Are requests claiming a Meta crawler identity verified before receiving WAF, rate-limit, or access exemptions?

Priority: 5
What
Evidence-backed tactic: Use User-Agent matching for classification without treating it as sufficient proof of trusted identity.
How
Match the documented token and use a Meta-published source-IP method only where the official page provides one. Do not borrow or invent ranges for newer agents.
Why
A request should not gain privileged treatment merely by presenting a documented User-Agent.

Can Meta crawler requests fetch public URLs without triggering writes, purchases, sessions, or other state changes?

Priority: 5
What
Evidence-backed tactic: Keep anonymously crawlable URLs read-only and safe to repeat.
How
Require authenticated intent, authorization, and confirmation for state-changing operations. Protect them against CSRF. Reject unsupported methods. Verify that anonymous GET and HEAD requests have no side effects.
Why
Meta says ExternalFetcher can support agentic site navigation for user tasks and may bypass robots.txt.

Are Meta crawler preferences expressed in robots.txt rather than relying on NoAI tags?

Priority: 5
What
Confirmed platform requirement: Encode Meta crawl choices through the relevant robots.txt agent groups.
How
Audit the production root robots file, add an explicit group for each applicable Meta token, validate the served response, and keep licensing notices separate from crawler enforcement.
Why
Meta identifies robots.txt as its supported industry-standard preference mechanism and contrasts it with nonstandard NoAI tags.

Is Applebot allowed where Apple search and source-linked answer visibility is desired?

Priority: 5
What
Confirmed platform requirement: Permit Apple's search crawler.
How
audit Applebot's robots group, HTTP access, and logs.
Why
Apple says crawled data powers Spotlight, Siri, Safari, and current-context AI answers.

Is Applebot-Extended policy independent from Applebot search access?

Priority: 5
What
Confirmed platform requirement: Choose foundation-model training use separately.
How
configure the Extended token without blocking Applebot if search visibility remains desired.
Why
Applebot-Extended is a use-control token and does not crawl pages itself.

Is `nosnippet` used only when content should be excluded from Apple AI answer context?

Priority: 5
What
Confirmed platform requirement: Control answer grounding while preserving ordinary discoverability.
How
apply Applebot-specific meta or header rules to exact content.
Why
Apple says `nosnippet` opts content out of broad-world-knowledge AI answers.

Is paywalled or metered content marked `isAccessibleForFree: false` for Applebot?

Priority: 5
What
Confirmed platform requirement: Declare page-level restricted access.
How
add accurate JSON-LD and validate it against the served page.
Why
Apple keeps marked pages eligible for results but excludes their content from AI answer context.

Are Applebot controls for PDFs, images, and other non-HTML resources sent through `X-Robots-Tag`?

Priority: 4
What
Confirmed platform requirement: Cover assets that cannot contain HTML meta tags.
How
inspect headers at CDN and origin for each file type.
Why
Apple documents header-level directives for non-HTML resources.

Does robots.txt explicitly mention Applebot rather than accidentally inheriting Googlebot rules?

Priority: 4
What
Confirmed platform requirement: Remove ambiguous fallback behavior.
How
add a deliberate Applebot group and test it.
Why
Apple says it follows Googlebot instructions when Applebot is absent but Googlebot is present.

Are Applebot search, answer-context, and model-training outcomes tested as three separate states?

Priority: 4
What
Evidence-backed tactic: Verify the intended combination.
How
check search discoverability, snippet/context directives, and Applebot-Extended policy independently.
Why
Apple's documented controls affect different product uses.

Is Applebot traffic authenticated using Apple's documented verification approach before WAF decisions?

Priority: 4
What
Confirmed platform requirement: Distinguish real traffic from spoofed UA strings.
How
follow Apple's bot-verification instructions and retain evidence in security logs.
Why
access decisions based on a bare UA are unsafe.

Do eligible pages return a real `200` response with substantive content?

Priority: 5
What
Confirmed platform requirement: Verify status and body together.
How
fetch from origin and edge as each target bot.
Why
A success code sends content to processing, but empty/error bodies can become soft 404s.

Are application error pages prevented from returning `200`?

Priority: 5
What
Confirmed platform requirement: Eliminate soft 404s.
How
map missing, expired, and failed records to meaningful statuses and test SPA fallbacks.
Why
Crawlers may treat error-like `200` content as absent.

Do removed resources return `404` or `410` rather than a thin success page?

Priority: 5
What
Confirmed platform requirement: Communicate permanent absence.
How
test deleted URLs through CDN, redirects, and application routes.
Why
Correct status accelerates removal and avoids stale retrieval.

Do permanent moves use a direct permanent redirect with a short chain?

Priority: 5
What
Confirmed platform requirement: Consolidate old URLs to the final canonical.
How
test every hop, status, and target response.
Why
Permanent redirects are canonical signals and long chains waste crawl work.

Are `429` and `5xx` responses rare, monitored, and correctly scoped?

Priority: 5
What
Confirmed platform requirement: Prevent crawl throttling and eventual index loss.
How
alert by bot, route, edge, and origin. Repair capacity rather than using errors as routine rate control.
Why
Google slows crawling on `429`/`5xx` and may eventually drop persistent failures.

Is `/robots.txt` highly available with intentional status behavior?

Priority: 5
What
Confirmed platform requirement: Treat it as critical infrastructure.
How
bypass brittle application dependencies, monitor content/status, and test failure modes.
Why
A robots `5xx` can stop Google crawling while it retries.

Are DNS, TLS, IPv4/IPv6, and edge routing healthy from external bot regions?

Priority: 5
What
Evidence-backed tactic: Verify the connection before HTML concerns.
How
synthetic-probe resolution, certificate chain, handshake, and response from multiple geographies.
Why
Network failures prevent any retrieval.

Are intended bots exempt from CAPTCHA, JavaScript challenge, interstitial, and login loops?

Priority: 5
What
Confirmed platform requirement: Remove non-content gates from crawler paths.
How
combine verified identity with narrowly scoped WAF rules and test clean sessions.
Why
OpenAI identifies these controls as common crawler failures.

Are country and ASN blocks checked against crawler egress locations?

Priority: 5
What
Evidence-backed tactic: Prevent invisible geo-denial.
How
test US and other documented crawler locations through the complete edge policy.
Why
Google primarily egresses from the US and may crawl from elsewhere.

Are `ETag`, `Last-Modified`, and correct `304` responses implemented for stable resources?

Priority: 3
What
Evidence-backed tactic: Enable efficient recrawling.
How
validate conditional requests and ensure updates invalidate validators.
Why
Google supports both cache validators and can reuse unchanged content.

Does critical text occur before crawler file-size cutoffs?

Priority: 4
What
Confirmed platform requirement: Keep primary content and metadata early and responses lean.
How
measure uncompressed HTML/resource sizes and test truncation.
Why
Google crawlers stop processing beyond product-specific limits.

Do APIs, media, CSS, and JavaScript return correct MIME types and unrestricted bot responses?

Priority: 4
What
Evidence-backed tactic: Ensure dependencies are fetchable and interpretable.
How
crawl resource graphs and compare status, content type, cache, and WAF results.
Why
Blocked or malformed resources can hide meaning and markup.

Are search, training, and user-triggered fetchers monitored as separate agents?

Priority: 5
What
Evidence-backed tactic: Preserve user agent, verified source, URL, status, bytes, and timing.
How
Create purpose-specific log views rather than one “AI bot” bucket.
Why
Access decisions and diagnostic meaning differ by crawler role.

Are claimed AI crawlers validated with official IP sources or documented verification rather than user agent alone?

Priority: 5
What
Evidence-backed tactic: Separate genuine platform access from spoofed bot traffic.
How
Match user agent and current official IP data, retaining verification status.
Why
User-agent strings are easy to spoof.

Are ChatGPT search inclusion and model-training preferences configured and audited independently?

Priority: 5
What
Confirmed platform requirement: Treat OAI-SearchBot and GPTBot as different controls.
How
Review robots rules and logs for each agent after changes.
Why
OpenAI states that search inclusion and training opt-out use separate bots.

Are `PerplexityBot` indexing requests distinguished from `Perplexity-User` user-triggered fetches?

Priority: 5
What
Confirmed platform requirement: Monitor and configure the two roles independently.
How
Use official user agents and current published IP ranges in WAF and logs.
Why
Perplexity documents different purposes and robots behavior for the two agents.

Are `ClaudeBot`, `Claude-User`, and other documented Anthropic agents configured according to intended use?

Priority: 4
What
Confirmed platform requirement: Avoid using one robots rule as a proxy for all Anthropic access.
How
Review the living crawler documentation before each policy change.
Why
Training collection and user-requested access are distinct purposes.

Are verified answer-engine requests monitored for `403`, challenge, rate-limit, and origin errors?

Priority: 5
What
Evidence-backed tactic: Detect accidental denial outside robots.txt.
How
Alert on status shifts by verified agent and sample affected URLs.
Why
A permissive robots file does not override a blocking WAF or CDN.

Are representative public URLs automatically checked for robots, indexability, renderability, and fetch status after deployments?

Priority: 4
What
Evidence-backed tactic: Detect access regressions before dashboard trends decline.
How
Test an English/Turkish canary set through public paths and verified log evidence.
Why
Small configuration changes can block an entire content class.

When using `noindex`, can the intended crawler fetch the page and read the directive?

Priority: 5
What
Confirmed platform requirement: Avoid blocking the crawler before it sees the exclusion signal.
How
Test robots access and rendered meta or header directives together.
Why
OpenAI notes that its crawler must access the page to detect `noindex`.

If Cloudflare Markdown for Agents is enabled, does the origin send an explicit `Content-Signal` header that matches the approved policy?

Priority: 4
What
Confirmed platform requirement: Prevent converted Markdown responses from silently receiving Cloudflare's permissive default.
How
Set the header at the origin, then request representative URLs with `Accept: text/markdown` and verify the effective response across routes and caches.
Why
Cloudflare preserves an origin header, but otherwise adds `ai-train=yes, search=yes, ai-input=yes` to converted Markdown.

Is the primary answer-bearing text present in the initial HTML?

Priority: 5
What
Evidence-backed tactic: Avoid making retrieval depend entirely on client execution.
How
compare raw response with rendered DOM for headings, facts, links, and citations.
Why
Google recommends server/pre-rendering and notes that not all bots run JavaScript.

Does Google's rendered HTML contain the same intended content and directives as the source response?

Priority: 5
What
Confirmed platform requirement: Detect hydration and rendering loss.
How
use URL Inspection rendered HTML and screenshot, then compare canonical, robots, structured data, and body.
Why
Google indexes rendered HTML for JavaScript pages.

Is primary content available without clicks, swipes, typing, or consent interaction?

Priority: 5
What
Confirmed platform requirement: Remove interaction-gated information from the retrieval path.
How
inspect clean-session HTML and rendered DOM before any action.
Why
Google does not trigger user interaction to lazy-load primary content.

Are render-critical JavaScript and CSS paths crawlable?

Priority: 5
What
Confirmed platform requirement: Prevent incomplete rendering.
How
crawl the resource graph and inspect robots rules, responses, and CSP/CDN failures.
Why
Google will not render JavaScript from blocked files or pages.

Are discoverable links rendered as real crawlable anchors with resolvable URLs?

Priority: 5
What
Confirmed platform requirement: Expose navigation and citations in standard HTML.
How
inspect raw/rendered anchors and crawl from hubs without a browser history API.
Why
Crawlers primarily discover URLs through links.

Does every indexable page declare one valid canonical matching the intended public URL?

Priority: 5
What
Confirmed platform requirement: Consolidate duplicates.
How
validate status, absolute URL, indexability, and reciprocal internal signals.
Why
Canonicalization decides which version can carry index and citation signals.

Do redirects, HTML/header canonicals, sitemap URLs, hreflang, and internal links agree?

Priority: 5
What
Evidence-backed tactic: Remove contradictory URL identity signals.
How
build a per-page signal matrix and fail conflicting targets.
Why
Consistency improves consolidation and crawl efficiency.

Are tracking, sort, filter, print, session, and case variants prevented from becoming competing answer candidates?

Priority: 4
What
Evidence-backed tactic: Control duplicate URL inventory.
How
normalize links, redirect true duplicates, canonicalize when needed, and block useless crawl spaces.
Why
Duplicate inventory consumes crawl resources and fragments identity.

Do canonical signals cover PDFs and other non-HTML documents?

Priority: 4
What
Confirmed platform requirement: Consolidate equivalent downloadable and HTML versions.
How
serve a `Link: <...>; rel="canonical"` response header where appropriate.
Why
HTML canonical tags cannot live inside non-HTML files.

Are mobile and desktop robots directives, title, description, canonicals, and structured data equivalent?

Priority: 5
What
Confirmed platform requirement: Avoid device-specific eligibility loss.
How
compare served and rendered variants.
Why
Google indexes the mobile version and warns that divergent directives can block indexing.

Does each citation-worthy document and media asset have a stable, directly fetchable URL?

Priority: 4
What
Evidence-backed tactic: Avoid transient blobs, expiring signed links, or app-only destinations.
How
test long-lived URLs without session state.
Why
Answer systems need resolvable sources to retrieve and link.

Is page structure encoded with semantic headings, landmarks, accessible names, roles, and states?

Priority: 3
What
Evidence-backed tactic: Make content and actions machine-readable.
How
run accessibility-tree and agent-task audits.
Why
OpenAI says Atlas uses ARIA to interpret interactive pages.

Are indexable routes free of hash-fragment-only content addressing?

Priority: 4
What
Confirmed platform requirement: Give content stable server-resolvable paths.
How
request every canonical URL without prior app state and migrate legacy AJAX fragments.
Why
Google deprecated the AJAX crawling scheme and does not rely on URL fragments for separate documents.

Is structured data present and valid in the rendered page after deployment?

Priority: 5
What
Confirmed platform requirement: Test the output crawlers receive alongside the templates.
How
run Rich Results Test, URL Inspection, and Bing markup inspection on production URLs.
Why
templating or serving failures can break valid source code after release.

Does the XML sitemap contain only preferred canonical, indexable production URLs?

Priority: 5
What
Confirmed platform requirement: Publish a clean discovery inventory.
How
diff sitemap entries against canonicals, status, noindex, environment, and redirects.
Why
Search engines use sitemaps as canonical and discovery hints.

Are sitemap locations fully qualified, absolute, encoded URLs that return the intended resource?

Priority: 5
What
Confirmed platform requirement: Eliminate malformed discovery targets.
How
parse, normalize, fetch, and compare each sampled URL.
Why
Google crawls URLs exactly as listed.

Are large sitemaps split within protocol limits and referenced by valid indexes?

Priority: 4
What
Confirmed platform requirement: Keep every file processable.
How
enforce at most 50 MB uncompressed or 50,000 URLs per sitemap and validate indexes.
Why
Oversized files can fail processing.

Does `lastmod` change only after a significant page update?

Priority: 5
What
Confirmed platform requirement: Send truthful freshness signals.
How
derive it from main content, structured data, or meaningful link changes. Do not use deploy or sitemap time.
Why
Bing uses accurate values to focus crawling.

Are sitemap `priority` and `changefreq` excluded from GEO scoring and operational promises?

Priority: 3
What
Confirmed platform requirement: Stop relying on ignored fields.
How
remove dashboards and checks that treat them as freshness/rank inputs.
Why
Bing says both fields are ignored.

Are sitemap submission status, last-read time, and processing errors monitored?

Priority: 5
What
Confirmed platform requirement: Confirm that discovery infrastructure is actually consumed.
How
alert on stale reads, fetch failures, and parse errors in Google and Bing webmaster tools.
Why
A published but unread sitemap provides no timely signal.

Where IndexNow is used, does the pipeline handle publish, substantive update, redirect, and deletion events without duplicate storms?

Priority: 2
What
Evidence-backed tactic: Operate optional change notifications as a reliable, idempotent freshness channel.
How
Queue canonical URLs, debounce cosmetic changes, retry from response codes, log receipts, and reconcile missed events.
Why
IndexNow recommends automated lifecycle notifications while discouraging repeat submissions without meaningful changes.

Are RSS/Atom feeds available and submitted for rapidly changing editorial content?

Priority: 4
What
Evidence-backed tactic: Provide a recent-change stream alongside the full sitemap.
How
validate entries, canonical links, timestamps, and feed fetchability.
Why
Google accepts RSS 2.0 and Atom 1.0 as sitemap formats.

Has the retired Google sitemap ping endpoint been removed?

Priority: 3
What
Confirmed platform requirement: Stop sending no-op HTTP pings.
How
submit through robots.txt, Search Console/API, sitemap fetch, RSS/Atom, or WebSub as appropriate.
Why
Google deprecated the ping endpoint and it returns `404`.

Do direct product feeds and public landing pages update atomically?

Priority: 4
What
Evidence-backed tactic: Prevent an answer engine from seeing mismatched availability, price, URL, or description.
How
publish with versioned jobs, validate both surfaces, then notify discovery systems.
Why
OpenAI notes delays and recommends direct feeds for current product data.

Is every GEO measurement tied to a specific decision rather than a generic visibility score?

Priority: 5
What
Evidence-backed tactic: State the decision, audience, metric, decision threshold, and owner.
How
Put these fields at the top of every recurring report and experiment brief.
Why
Metrics without a decision invite post-hoc narratives and metric shopping.

Does each GEO metric have a named owner, review cadence, and escalation recipient?

Priority: 4
What
Evidence-backed tactic: Record accountability for collection, interpretation, and remediation.
How
Maintain a lightweight RACI beside the metric dictionary.
Why
Unowned dashboards decay and alerts go unanswered.

Are answer mentions, citations, referral visits, conversions, and revenue reported as separate outcome layers?

Priority: 5
What
Evidence-backed tactic: Use a funnel with distinct denominators instead of one blended GEO score.
How
Define a metric for each observable layer and disclose missing links between them.
Why
A citation is neither a click nor a conversion.

Are GEO metrics defined with numerator, denominator, unit, scope, and known blind spots?

Priority: 4
What
Evidence-backed tactic: Document exact formulas for prevalence, citation share, accuracy, and conversion metrics.
How
Version a shared metric dictionary and link it from every dashboard.
Why
Similar labels often hide incompatible calculations.

Is the exact set of answer engines, surfaces, accounts, and markets in scope recorded?

Priority: 5
What
Evidence-backed tactic: Name each product surface rather than grouping everything under “AI search.”
How
Track platform, surface, access mode, subscription state, language, and country.
Why
Products can retrieve, cite, personalize, and report data differently.

Does monitoring include the brand's official name, spelling variants, products, people, and common aliases?

Priority: 5
What
Evidence-backed tactic: Maintain the entity strings that count as a mention and those that create ambiguity.
How
Review aliases with brand, legal, and local-language owners quarterly.
Why
Exact-string matching undercounts variants and can merge unrelated entities.

Is the owned, operated, partner, reseller, and unaffiliated domain set explicitly classified?

Priority: 4
What
Evidence-backed tactic: Establish which cited hosts count as first-party visibility.
How
Keep a versioned registry including subdomains, migrations, and country domains.
Why
Host-level aggregation can misattribute partner or legacy properties.

Is the competitor or peer set versioned with clear inclusion rules?

Priority: 3
What
Evidence-backed tactic: Preserve the entities used in share and recommendation comparisons.
How
Add or remove peers only at declared reporting boundaries.
Why
Changing the peer set changes normalized shares even if engine behavior is unchanged.

Are start time, end time, timezone, cadence, and blackout periods recorded?

Priority: 5
What
Evidence-backed tactic: Make the temporal sampling frame explicit.
How
Store UTC timestamps and a business timezone for every run.
Why
Engines and indexed content change over time.

Is each observation accompanied by platform, surface, prompt, context, locale, account state, and timestamp?

Priority: 5
What
Evidence-backed tactic: Store enough metadata to explain or reproduce a run.
How
Use an immutable run ID and structured observation schema.
Why
A response without execution context is weak evidence.

Are full response text, citations, visible source panels, and screenshots retained subject to policy?

Priority: 5
What
Evidence-backed tactic: Keep auditable evidence behind derived scores.
How
Store a timestamped snapshot with access controls and retention limits.
Why
Interfaces and citations can change after collection.

Are direct brand and navigational prompts measured separately from discovery prompts?

Priority: 4
What
Evidence-backed tactic: Create a dedicated segment for users already seeking the entity.
How
Tag brand-only, brand-plus-topic, and URL/navigation prompts.
Why
High visibility on navigational queries can mask weak category discovery.

Does the prompt set include unbranded category and problem-discovery needs?

Priority: 5
What
Evidence-backed tactic: Measure whether the brand appears before it is named.
How
Source prompts from search demand, support logs, sales calls, and user research.
Why
Unbranded discovery is a different exposure opportunity from brand recall.

Are “best,” “alternative,” “versus,” and shortlist prompts analyzed as a distinct segment?

Priority: 5
What
Evidence-backed tactic: Track inclusion, exclusion, stated criteria, and cited evidence in comparative answers.
How
Build balanced prompts across competitors and decision criteria.
Why
Comparative answers carry higher commercial and brand-risk stakes.

Are informational prompts that should cite first-party facts measured separately from recommendation prompts?

Priority: 4
What
Evidence-backed tactic: Identify questions about specifications, policies, people, prices, and processes.
How
Map each prompt to a maintained source-of-truth URL and expected fact set.
Why
Factual accuracy and commercial recommendation require different rubrics.

Is each prompt tagged by awareness, consideration, decision, support, or retention intent?

Priority: 3
What
Evidence-backed tactic: Connect visibility to the expected next user action.
How
Apply a documented intent rubric and double-review ambiguous prompts.
Why
Equal weighting across unlike intents distorts business significance.

Are English and Turkish prompt results collected and reported as separate cohorts?

Priority: 5
What
Evidence-backed tactic: Preserve language-specific exposure, accuracy, and citation metrics.
How
Use native prompts rather than literal translations and compare only equivalent intents.
Why
Aggregation can hide language-specific failure and source selection.

Are country, interface locale, and query language recorded independently?

Priority: 4
What
Evidence-backed tactic: Distinguish geography from language in every observation.
How
Record declared location, observed location, locale, timezone, and VPN/proxy use.
Why
Search results and query rewrites can vary with general location.

Is logged-in, memory-enabled, or personalized testing separated from clean-session testing?

Priority: 4
What
Evidence-backed tactic: Treat user context as an experimental factor.
How
Record account state and run a stable clean-session benchmark alongside real-user scenarios.
Why
Memory and conversation context can alter query rewriting and responses.

Are desktop, mobile, app, browser, and embedded surfaces identified in observations?

Priority: 3
What
Evidence-backed tactic: Preserve the interface through which each answer was generated.
How
Record device class and product surface with each run.
Why
Available presentation and local features can vary by surface.

Is the benchmark grounded in real user needs rather than only synthetic prompts?

Priority: 5
What
Evidence-backed tactic: Blend consented production questions, research, search demand, and expert-authored edge cases.
How
Record provenance and sampling rules for every prompt.
Why
A convenient prompt list can measure the benchmark designer rather than the audience.

Does the reporting retain real prompt-frequency weights while also showing an unweighted diagnostic view?

Priority: 4
What
Evidence-backed tactic: Separate business-weighted performance from coverage across unique needs.
How
Store both the occurrence weight and one-vote-per-prompt result.
Why
Deduplicating all repeated needs can erase demand concentration.

Does the corpus include ambiguous, misspelled, adversarial, and high-risk fact prompts?

Priority: 4
What
Evidence-backed tactic: Test realistic failure modes beyond the happy path.
How
Maintain an edge-case suite informed by incidents, support tickets, and red-team review.
Why
Average visibility can coexist with severe brand-safety failures.

Are personal, confidential, contractual, and embargoed details removed before prompts reach external systems?

Priority: 5
What
Evidence-backed tactic: Apply data minimization and approval rules to the test corpus.
How
Redact identifiers, use synthetic substitutes, and prohibit secret-bearing prompts.
Why
GEO monitoring should not become an uncontrolled disclosure channel.

Is every benchmark run tied to an immutable prompt-set version?

Priority: 5
What
Evidence-backed tactic: Preserve prompt text, metadata, weights, additions, removals, and rationale.
How
Use a versioned repository or append-only dataset manifest.
Why
Trend lines are uninterpretable when the denominator silently changes.

Is a stable subset of prompts retained for longitudinal comparison?

Priority: 4
What
Evidence-backed tactic: Keep a frozen panel while allowing a separate evolving discovery corpus.
How
Version both and never splice their trend lines without restatement.
Why
Corpus refreshes otherwise masquerade as performance change.

Is an unseen prompt subset reserved to test whether an optimization generalizes?

Priority: 4
What
Evidence-backed tactic: Separate tuning prompts from validation prompts.
How
Restrict access to the holdout and evaluate it only at planned gates.
Why
Repeatedly optimizing on the same prompts overfits the benchmark.

Are inclusion, exclusion, deduplication, and retirement rules documented?

Priority: 3
What
Evidence-backed tactic: Make corpus curation reproducible.
How
Require a reason code and reviewer for every material corpus change.
Why
Selective prompt removal can manufacture an improvement.

Does each factual prompt identify the authoritative owned page or record expected to support it?

Priority: 5
What
Evidence-backed tactic: Create a prompt-to-evidence mapping.
How
Store expected URLs, approved facts, freshness owner, and review date.
Why
A brand cannot assess answer correctness without a maintained reference point.

Are accuracy, relevance, tone, citation, and safety criteria specified before scoring?

Priority: 4
What
Evidence-backed tactic: Use a multi-dimensional rubric rather than a single “good answer” judgment.
How
Define anchored scales and unacceptable-failure conditions for each prompt class.
Why
Vague criteria produce inconsistent review and retrospective scoring.

Is the cited page content captured or hashed at observation time?

Priority: 4
What
Evidence-backed tactic: Preserve evidence of what the engine could have retrieved then.
How
Store a permitted snapshot or content checksum plus fetch timestamp and status.
Why
Later page edits can be mistaken for engine inconsistency.

Are timeouts, captchas, no-search responses, parsing failures, and partial answers retained as outcomes?

Priority: 4
What
Evidence-backed tactic: Make missingness visible instead of silently dropping failed runs.
How
Use explicit failure reason codes and report the failure rate.
Why
Non-random missing observations can bias visibility upward.

Does measurement distinguish answers that searched the web from answers that did not?

Priority: 4
What
Evidence-backed tactic: Record observable search activation and citation availability.
How
Use interface indicators where available and label uncertain cases as unknown.
Why
Absence of a citation is not comparable when retrieval never occurred.

Are raw observations retained alongside deduplicated and normalized metrics?

Priority: 4
What
Evidence-backed tactic: Preserve citation events before host, page, or entity aggregation.
How
Build transformations from immutable raw records with versioned code.
Why
Aggregation rules can change conclusions and must be auditable.

Are browser, collector, extraction, classifier, and rubric versions attached to every run?

Priority: 4
What
Evidence-backed tactic: Treat measurement code changes as potential discontinuities.
How
Store commit or release identifiers and annotate deployments.
Why
A parser change can look like a visibility change.

Is every important prompt sampled repeatedly rather than treated as a single rank check?

Priority: 5
What
Evidence-backed tactic: Estimate the response distribution for the same prompt and context.
How
Repeat at precommitted times and report the number of valid runs.
Why
Identical prompts can return different answers and citations.

Are observations spread across multiple days and relevant dayparts?

Priority: 5
What
Evidence-backed tactic: Capture temporal variability rather than a single collection burst.
How
Use a scheduled panel with timestamps and a stable prompt mix.
Why
Same-session bursts can understate non-stationarity.

Is prompt execution order randomized or counterbalanced?

Priority: 3
What
Evidence-backed tactic: Prevent time, throttling, or session-order effects from aligning with one treatment.
How
Randomize within blocks of platform, locale, and intent.
Why
Fixed order can confound treatment with collection conditions.

Is sample size chosen before looking at favorable results?

Priority: 5
What
Evidence-backed tactic: Base collection volume on pilot variance, desired precision, and cost.
How
Write the stopping rule in the study plan and keep it fixed.
Why
Optional stopping inflates false discoveries.

Does collection continue to the precommitted endpoint even after a favorable result appears?

Priority: 5
What
Evidence-backed tactic: Enforce the planned stopping rule.
How
Lock the schedule or require documented independent approval to stop early.
Why
Stopping on a lucky result manufactures confidence.

Are visibility estimates reported with uncertainty intervals rather than point estimates alone?

Priority: 5
What
Evidence-backed tactic: Pair prevalence, share, and quality metrics with interval estimates.
How
Use an appropriate bootstrap or model and disclose assumptions.
Why
Small apparent differences may be measurement noise.

Does every score show prompts attempted, valid responses, and observations contributing to it?

Priority: 5
What
Evidence-backed tactic: Expose denominators and missing data.
How
Display `n` beside each segment and suppress underpowered comparisons.
Why
A percentage without its denominator looks more certain than it is.

Are low-frequency domains and brands flagged as especially unstable?

Priority: 4
What
Evidence-backed tactic: Treat sparse citation prevalence cautiously.
How
Show zero counts, intervals, and a low-base warning rather than forcing ranks.
Why
Rare events produce volatile percentages and ordinal positions.

Is platform drift tested before combining old and new observations?

Priority: 4
What
Evidence-backed tactic: Look for level, variance, or source-distribution changes over time.
How
Maintain rolling diagnostics and annotate platform or collector changes.
Why
Historical samples may no longer estimate the current system.

Does the methodology avoid declaring one fixed run count sufficient for every platform and prompt class?

Priority: 4
What
Evidence-backed tactic: Calibrate precision by local variance and decision risk.
How
Run a pilot, estimate uncertainty, then set a documented collection plan.
Why
Variability differs by platform, topic, language, and outcome.

Are treatment comparisons matched on platform, prompt, locale, surface, and collection window?

Priority: 5
What
Evidence-backed tactic: Block or pair observations on major context variables.
How
Use the same sampling schedule and prompt version for control and treatment.
Why
Cross-context differences can exceed the treatment effect.

Are platform-specific metrics shown before any blended index?

Priority: 5
What
Evidence-backed tactic: Preserve each platform's distinct citation and retrieval behavior.
How
If a composite is used, disclose weights and show components.
Why
Raw citation counts are not directly comparable across platforms.

Is citation share normalized within a declared platform, prompt set, and peer universe?

Priority: 4
What
Evidence-backed tactic: State exactly what total forms the denominator.
How
Publish both raw citation count and normalized share.
Why
“Share of voice” is otherwise easy to misread as market share.

Do major conclusions hold under reasonable alternative prompt and business weights?

Priority: 3
What
Experimental practice: Test whether one weighting scheme drives the result.
How
Report weighted, unweighted, and segment-level views.
Why
Hidden weights can reverse comparative rankings.

Is the smallest decision-relevant improvement defined before an experiment?

Priority: 4
What
Evidence-backed tactic: Distinguish practical impact from any nonzero numerical change.
How
Set an effect threshold with stakeholders and power the study accordingly.
Why
Tiny changes can be statistically noisy or commercially irrelevant.

Is brand mention prevalence measured as the share of valid responses containing the entity?

Priority: 5
What
Evidence-backed tactic: Distinguish response-level presence from citation count.
How
Apply the versioned alias set and publish the valid-response denominator.
Why
Repeated citations in one answer should not inflate how often the brand appears.

Is owned-domain citation prevalence tracked separately from brand mentions?

Priority: 5
What
Evidence-backed tactic: Count responses with at least one owned citation.
How
Resolve cited URLs against the versioned owned-domain registry.
Why
Engines can mention a brand without citing it, or cite it without naming it prominently.

Is raw citation count retained as a descriptive metric without being treated as a rank?

Priority: 3
What
Evidence-backed tactic: Count observed citation events before deduplication.
How
Report count beside prevalence, share, and sample size.
Why
Counts are useful operationally but depend heavily on response and platform format.

Is the number and distribution of unique owned pages cited tracked?

Priority: 4
What
Evidence-backed tactic: Measure whether visibility depends on one URL or a resilient content set.
How
Canonicalize URLs and report concentration by page.
Why
A domain total can hide fragile dependence on a single asset.

Is the diversity and concentration of all cited source domains monitored?

Priority: 3
What
Evidence-backed tactic: Track how concentrated answer evidence is across publishers.
How
Report unique domains and a concentration statistic by prompt segment.
Why
Source concentration can expose dependency and information-quality risks.

Are cited owned pages classified as product, service, research, guide, profile, policy, or support content?

Priority: 3
What
Experimental practice: Identify which asset types engines use for each intent.
How
Map canonical URLs to a stable content taxonomy.
Why
Page-type patterns can guide testing without pretending to reveal a ranking factor.

Is the entity's narrative prominence scored separately from citation occurrence?

Priority: 3
What
Experimental practice: Assess whether the brand is central, supporting, incidental, or absent.
How
Use an anchored human rubric and preserve the answer text.
Why
A footnote-level source and a central recommendation create different exposure.

Does reporting avoid treating Bing citation totals or URL counts as placement data?

Priority: 5
What
Confirmed platform requirement: Label placement as unavailable unless directly observed and preserved.
How
Remove rank-like labels from Bing AI Performance exports.
Why
The platform explicitly says these metrics do not show placement or importance.

Is positive, neutral, negative, mixed, and purely factual treatment scored with an anchored rubric?

Priority: 4
What
Experimental practice: Separate visibility from the answer's stance toward the entity.
How
Use human-reviewed examples and an explicit “not applicable” state.
Why
More mentions can be harmful when the context is inaccurate or adverse.

Is shortlist or recommendation inclusion measured only for prompts where a recommendation is appropriate?

Priority: 5
What
Experimental practice: Record included, excluded, warned against, or not applicable.
How
Apply intent-specific scoring and retain the stated selection criteria.
Why
Treating every mention as a recommendation overstates commercial exposure.

Is the cited page directly relevant to the user's question and cited claim?

Priority: 4
What
Evidence-backed tactic: Distinguish topical overlap from evidence that answers the claim.
How
Use an anchored relevance scale with examples.
Why
Broadly related sources can create a false appearance of support.

Are brand, product, people, policy, location, and pricing facts checked against current authoritative records?

Priority: 5
What
Evidence-backed tactic: Score fact-level correctness and materiality.
How
Maintain a dated truth set and require subject-matter review for high-risk claims.
Why
Citation presence does not guarantee the answer is correct.

Are responses checked for merging the brand with similarly named organizations, people, or products?

Priority: 5
What
Evidence-backed tactic: Record identity collisions as a distinct critical error.
How
Test aliases and ambiguous prompts. Verify official identifiers and domains.
Why
An answer can appear visible while describing the wrong entity.

Are claims credited to the correct organization, author, study, or page?

Priority: 5
What
Evidence-backed tactic: Score attribution independently from factual truth.
How
Compare the answer's attribution with the cited source's authorship and context.
Why
Correct facts can still damage trust when assigned to the wrong source.

Are cited facts and pages checked for current validity at collection time?

Priority: 5
What
Evidence-backed tactic: Track source date, last verified date, and whether the claim is time-sensitive.
How
Apply shorter review windows to pricing, people, policies, and regulatory facts.
Why
A well-cited answer can still be stale.

Is the share of citations to primary, official, or original evidence tracked for high-risk prompts?

Priority: 4
What
Evidence-backed tactic: Classify sources by provenance rather than domain popularity alone.
How
Use a documented source-tier rubric and review exceptions.
Why
Secondary summaries can introduce drift or strip conditions from claims.

Are high-risk answers reviewed for reliance on apparently synthetic or low-accountability sources?

Priority: 3
What
Experimental practice: Examine source provenance, authorship, evidence chain, and editorial accountability.
How
Manually audit a risk-based sample and record uncertainty.
Why
Synthetic sources can amplify unsupported claims through recitation.

Is no source labeled AI-generated solely from an AI-content detector score?

Priority: 5
What
Evidence-backed tactic: Treat detector output as a weak triage signal without presenting it as provenance evidence.
How
Require corroborating authorship, disclosure, provenance, or editorial evidence.
Why
Detector accuracy is dataset- and model-dependent.

Do cited URLs resolve without errors, unsafe redirects, login walls, or removed content?

Priority: 4
What
Evidence-backed tactic: Measure whether users can inspect the cited evidence.
How
Fetch links at observation time and classify status, redirect, and access barrier.
Why
An inaccessible citation cannot provide practical verifiability.

Are URL parameters, fragments, redirects, and duplicate paths resolved before page-level reporting?

Priority: 4
What
Evidence-backed tactic: Map observed links to stable canonical identities while retaining raw URLs.
How
Follow safe redirects and use declared canonicals with anomaly checks.
Why
URL variants can inflate cited-page diversity.

Is repeated citation of the same page within one answer counted consistently?

Priority: 3
What
Evidence-backed tactic: Preserve events but define response-level deduplication.
How
Publish raw-event, unique-page, and response-prevalence metrics.
Why
Different interfaces repeat links differently.

Is inter-rater agreement measured for subjective GEO rubrics?

Priority: 4
What
Evidence-backed tactic: Quantify consistency for prominence, relevance, stance, and support judgments.
How
Double-score a stratified sample and adjudicate systematic disagreements.
Why
A human rubric is not reliable merely because humans applied it.

Are legal, safety, identity, pricing, and reputational failures confirmed by a qualified human?

Priority: 5
What
Evidence-backed tactic: Prevent automated scores from closing high-impact incidents.
How
Route critical flags to subject-matter, legal, or brand review.
Why
Automated graders can be confidently wrong.

Is every automated LLM grader compared with blinded human judgments on representative cases?

Priority: 5
What
Evidence-backed tactic: Validate judge accuracy before scaling its scores.
How
Maintain a gold set, confusion analysis, and periodic recalibration.
Why
An unvalidated judge can automate bias and drift.

Are automated judges asked for classification, pairwise choice, or anchored scores rather than open-ended impressions?

Priority: 4
What
Evidence-backed tactic: Constrain grading to auditable decisions.
How
Use explicit labels, evidence requirements, and an abstain option.
Why
Bounded tasks are easier to calibrate and reproduce.

Is every score tied to a rubric version and example set?

Priority: 4
What
Evidence-backed tactic: Treat criteria changes as measurement changes.
How
Store version, effective date, approver, and backfill policy.
Why
Quiet rubric drift corrupts longitudinal trends.

Are uncertain and disputed judgments retained instead of forced into a single label?

Priority: 3
What
Evidence-backed tactic: Add abstain, unclear, and adjudicated states.
How
Report disagreement rates and store both original ratings.
Why
Disagreement often reveals ambiguous prompts or weak criteria.

Are Turkish answers reviewed by fluent Turkish evaluators and English answers by fluent English evaluators?

Priority: 5
What
Evidence-backed tactic: Evaluate meaning, fluency, cultural nuance, and factual terminology in-language.
How
Use native or professionally fluent reviewers with a shared cross-language rubric.
Why
Machine translation can hide tone, ambiguity, and factual errors.

Are English/Turkish comparisons based on equivalent user intent rather than literal string translation?

Priority: 4
What
Evidence-backed tactic: Create paired prompts that are natural in each market.
How
Use forward translation, native rewriting, and independent intent review.
Why
Literal translation can change frequency, politeness, specificity, and retrieval behavior.

Are accuracy, citation support, visibility, and incident rates disaggregated by English and Turkish?

Priority: 5
What
Evidence-backed tactic: Make material parity gaps visible.
How
Show language-level intervals and prohibit a blended score from hiding failure.
Why
Overall averages can conceal poor performance for one language.

Do English and Turkish factual prompts map to equally current, authoritative first-party sources?

Priority: 4
What
Evidence-backed tactic: Audit missing translations and stale localized evidence.
How
Compare source-of-truth coverage and review dates across language pairs.
Why
A weaker Turkish evidence base can produce poorer citations and facts.

Does analytics preserve and classify referral URLs carrying `utm_source=chatgpt.com`?

Priority: 5
What
Confirmed platform requirement: Create a documented ChatGPT referral channel rule.
How
Test landing-page collection, redirects, consent, and warehouse ingestion end to end.
Why
OpenAI automatically adds this source tag to clicked referral URLs.

Are referral headers and campaign parameters preserved through the landing stack where privacy policy permits?

Priority: 4
What
Experimental practice: Prevent avoidable source loss between click and analytics event.
How
Test CDN, shortener, redirect, consent, and cross-domain paths.
Why
Redirect and parameter loss can misclassify visits.

Does reporting avoid attributing unexplained Direct traffic growth to answer engines?

Priority: 5
What
Evidence-backed tactic: Keep Direct as unknown-source traffic unless corroborated.
How
Use tagged referrals, experiments, surveys, and assisted-journey evidence.
Why
Many unrelated mechanisms create Direct/none sessions.

Are internal and external redirects tested for UTM and referrer preservation?

Priority: 4
What
Evidence-backed tactic: Detect source loss on the exact AI-referral landing paths.
How
Run browser tests and inspect analytics events after every redirect hop.
Why
A valid tagged click can arrive as Direct after faulty routing.

Are first-user, session, and event-scoped traffic-source dimensions used for their intended questions?

Priority: 5
What
Confirmed platform requirement: Avoid mixing acquisition, visit, and attributed-conversion scope.
How
Label scope in every chart and semantic model.
Why
GA4 applies different traffic-source scopes and attribution behavior.

Are answer-engine referrals evaluated beyond last-click conversion?

Priority: 5
What
Evidence-backed tactic: Analyze new-user acquisition, return visits, branded search, and downstream conversion paths.
How
Use consented user/session paths and a stated attribution model.
Why
Research-oriented visits may influence a later channel.

Are platform impressions, clicked referrals, and inferred “dark” exposure labeled separately?

Priority: 5
What
Evidence-backed tactic: Attach an evidence class to every KPI.
How
Use `observed platform`, `observed click`, `modeled`, or `unknown` labels.
Why
Combining them hides major identifiability limits.

Are AI-referred landing pages analyzed by intent, content type, language, and device?

Priority: 4
What
Evidence-backed tactic: Diagnose which experiences convert or fail after the click.
How
Build cohorts from validated source tags and canonical landing URLs.
Why
Channel-wide averages can mask high- and low-value entry points.

Are conversion rates based on qualified sessions or users with bot and internal traffic removed?

Priority: 5
What
Evidence-backed tactic: Define eligible traffic and conversion events before comparison.
How
Publish numerator, denominator, exclusions, and confidence interval.
Why
Tiny or contaminated denominators create dramatic but unstable rates.

Are macro conversions, micro conversions, duplicates, tests, and invalid events classified and audited?

Priority: 5
What
Evidence-backed tactic: Maintain a conversion-event dictionary with business meaning and firing rules.
How
Test events end to end and distinguish task completion from engagement proxies.
Why
Duplicate or soft events can manufacture apparent channel performance.

Are revenue per session, qualified-lead rate, deal quality, or equivalent value metrics compared by source?

Priority: 5
What
Evidence-backed tactic: Evaluate outcome quality alongside whether a conversion occurred.
How
Define a business-accepted value metric and apply minimum-sample safeguards.
Why
Channels can produce the same conversion rate with very different value.

Are engagement signals tied to the landing page's purpose rather than a generic session-duration target?

Priority: 3
What
Evidence-backed tactic: Choose actions such as reading, tool use, download, contact, or task completion.
How
Validate event instrumentation and segment by source and intent.
Why
A short visit may represent success when the answer is immediately found.

Are phone, meeting, proposal, retail, or other offline outcomes joined where lawful and useful?

Priority: 4
What
Evidence-backed tactic: Extend measurement beyond browser conversion events.
How
Use consented identifiers, controlled lookback rules, and documented match rates.
Why
High-consideration GEO influence may surface after the web session.

Are crawler requests, monitors, employees, agencies, and QA traffic excluded from human referral KPIs?

Priority: 5
What
Evidence-backed tactic: Maintain separate human analytics and crawler observability datasets.
How
Use validated IP, user-agent, network, and test identifiers without relying on one signal alone.
Why
GEO monitoring itself can contaminate the traffic it measures.

Are consent-denied and unmeasured visits acknowledged in GEO traffic reporting?

Priority: 5
What
Evidence-backed tactic: Separate observed analytics from total traffic claims.
How
Report consent rates, modeled fields, and market-specific collection differences.
Why
Measurement coverage can vary by jurisdiction, device, and browser.

Is the site's Referrer-Policy tested for unintended loss of useful, permitted referral context?

Priority: 3
What
Evidence-backed tactic: Understand what referrer data browsers send across origins.
How
Inspect headers on AI-referral landing pages and run cross-origin browser tests.
Why
Policy and link attributes can suppress the `Referer` header.

Is answer exposure without a click reported separately from referral traffic?

Priority: 5
What
Evidence-backed tactic: Use platform impression or sampled-answer evidence as leading indicators only.
How
Maintain distinct dashboards and avoid estimated click-through without data.
Why
A visible citation may satisfy the user without a site visit.

Does reporting avoid calculating click-through rate from a platform report that only provides impressions or citations?

Priority: 5
What
Confirmed platform requirement: Require a compatible click numerator and impression denominator.
How
Disable synthetic CTR fields for current Google Generative AI and Bing citation reports unless the platform adds documented clicks.
Why
Mixing unrelated datasets creates a meaningless ratio.

Are AI-referral conversion comparisons stratified by major traffic-mix factors and accompanied by raw sample sizes?

Priority: 4
What
Evidence-backed tactic: Compare like with like before attributing conversion differences to the referral channel.
How
Predefine relevant strata such as intent, landing page, market, device, and new-versus-returning status. Report each stratum's numerator, denominator, and uncertainty, or fit an adjusted model.
Why
Different traffic composition can confound aggregate channel comparisons. Stratification can reduce systematic comparison error.

Do prompt, response, analytics, and CRM joins have approved retention periods and role-based access?

Priority: 5
What
Evidence-backed tactic: Minimize exposure of user and commercial data.
How
Document purpose, fields, access roles, deletion, and audit logging.
Why
A richer GEO dataset creates a larger privacy and security surface.

Is the “up to 40% GEO lift” claim confined to its actual experiment?

Priority: 1
What
Experimental practice: Describe it only as a relative maximum on the study's visibility metric after five sources were already supplied in a fixed context.
How
Cite the original paper with the July 2026 critical survey. State that the setup was post-retrieval and did not test stable crawling, indexing, cross-platform discoverability, traffic, or business outcomes.
Why
The figure is widely misreported as a general organic-visibility or click lift that the evidence does not establish.

Are added citations, quotations, and statistics tested only when truthful and useful?

Priority: 1
What
Experimental practice: Evaluate whether extractable evidence improves source use without degrading retrieval or readability.
How
A/B test verified additions against an untreated baseline and audit factual support.
Why
The KDD experiment found post-retrieval gains for these tactics.

Has “authoritative tone” been rejected as a substitute for authority?

Priority: 1
What
Experimental practice: Do not make prose more forceful merely to influence a model.
How
improve evidence, expertise, qualifications, and uncertainty instead. Test tone only for reader clarity.
Why
The foundational paper found no significant general improvement from authoritative styling.

Has the team rejected a universal tiny-chunk or ideal-length rule?

Priority: 4
What
Evidence-backed tactic: Structure content around reader needs and coherent topics.
How
require evidence before enforcing paragraph, section, or page-length thresholds.
Why
Google explicitly says tiny “chunking” is not required and there is no ideal page length.

Is `llms.txt` excluded from Google visibility requirements and promises?

Priority: 4
What
Evidence-backed tactic: Treat it as an optional file only for consumers that explicitly document support.
How
remove Google ranking claims and do not divert maintenance from indexable content.
Why
Google says it ignores `llms.txt` for Search and generative features.

Has AI-only rewriting and exact prompt-variant publishing been rejected?

Priority: 4
What
Evidence-backed tactic: Write naturally for people and consolidate equivalent phrasings.
How
review briefs for long-tail variant pages, synonym stuffing, or a synthetic “AI voice.”
Why
Google says no special AI writing style or every-query variation is needed.

Are fabricated reviews, paid undisclosed mentions, and coordinated fake citations prohibited?

Priority: 5
What
Evidence-backed tactic: Reject inauthentic corroboration.
How
audit outreach, affiliate, influencer, review, and digital-PR programs for identity, disclosure, editorial independence, and spam.
Why
Google says seeking inauthentic mentions is unhelpful and spam systems still apply.

Has special “GEO schema” or guaranteed citation markup been rejected?

Priority: 4
What
Evidence-backed tactic: Use structured data for truthful semantics and documented features.
How
require an official consumer specification before adding a property and never promise AI inclusion.
Why
Google says structured data is not required for generative AI and no special schema exists.

Does every GEO test distinguish post-retrieval effects from organic discoverability?

Priority: 2
What
Experimental practice: Name whether the document was injected, retrieved, reranked, cited, absorbed, clicked, or converted.
How
prespecify the causal stage and denominator before data collection.
Why
Fixed-context studies cannot establish crawling or retrieval effects.

Are rewrites checked for upstream retrieval harm before rollout?

Priority: 2
What
Experimental practice: Ensure a passage optimized for citation has not lost query relevance or retrievability.
How
compare index/retrieval/rerank/citation outcomes and reader quality against the original.
Why
The survey reports an end-to-end benchmark where body-only optimization reduced upstream and final outcomes.

Are content experiments repeated across named engines and modes?

Priority: 2
What
Experimental practice: Avoid generalizing one answer surface to “AI search.”
How
record product, mode, version/date, account state, locale, and the exact prompts for each run.
Why
Commercial engines differ and vary over time.

Do experiments include paraphrases, repeated runs, time windows, baselines, and null outputs?

Priority: 2
What
Experimental practice: Estimate variability rather than showcase a favorable screenshot.
How
use multiple natural paraphrases, close repetitions, later reruns, untreated/placebo controls, and retain no-search/no-citation/error outcomes.
Why
The survey recommends these controls for reproducible GEO measurement.

Are citation counts paired with support, attribution, tone, and factual-use audits?

Priority: 2
What
Experimental practice: Measure whether the answer used the source correctly.
How
sample cited claims and rate entailment, accuracy, polarity, prominence, and missing qualifications with human review.
Why
Citation presence can coexist with unsupported or negative use.

Are citations separated from clicks, conversions, and business value?

Priority: 2
What
Experimental practice: Treat visibility as a vector rather than one rank.
How
report discoverability, citation, prominence, factual absorption, referral, assisted outcome, and conversion separately.
Why
Microsoft says citation counts do not indicate ranking or placement. Research finds no stable downstream causal result.

Does source-quality monitoring include concentration, synthetic-source risk, and diversity?

Priority: 2
What
Experimental practice: Audit what kinds of sources answer engines cite around priority topics.
How
classify ownership, originality, credibility, AI-generation evidence, geography, and repeated domain concentration. Manually verify flags.
Why
Recent audits find concentrated and sometimes synthetic source ecosystems.

Does every GEO change have a prewritten hypothesis, target segment, metric, and expected direction?

Priority: 5
What
Evidence-backed tactic: Define what result would support or refute the change.
How
Register the hypothesis before publishing the treatment.
Why
Post-hoc stories make every outcome look successful.

Is the main content or technical change distinguishable from unrelated edits?

Priority: 5
What
Evidence-backed tactic: Limit simultaneous variables or use a factorial design.
How
Keep a change manifest and postpone unrelated updates on test units.
Why
Bundled changes prevent causal learning.

Does the experiment include a comparable untreated control during the same period?

Priority: 5
What
Evidence-backed tactic: Separate treatment effects from platform-wide movement.
How
Match or randomize eligible pages, topics, markets, or prompt clusters.
Why
Before/after change alone is confounded by time.

Is randomization performed at a unit that avoids treatment contamination?

Priority: 4
What
Evidence-backed tactic: Choose page, topic cluster, locale, or market deliberately.
How
Document the unit, blocking variables, and spillover risks.
Why
Randomizing individual URLs can fail when engines synthesize across a whole site.

Are treatment and control pages comparable in intent, baseline demand, authority, freshness, and language?

Priority: 4
What
Evidence-backed tactic: Reduce baseline imbalance before comparison.
How
Pair or stratify units using pre-period data and domain expertise.
Why
Page selection bias can dominate the treatment.

Do users and Googlebot receive the same experiment logic and content eligibility?

Priority: 5
What
Evidence-backed tactic: Avoid bot-specific variants.
How
Use normal client/server experimentation without user-agent targeting.
Why
Google identifies test-page cloaking as a spam-policy violation.

Do alternate test URLs point to the preferred original with `rel="canonical"` where appropriate?

Priority: 5
What
Confirmed platform requirement: Signal the preferred URL during multi-URL tests.
How
Validate rendered tags and ensure the canonical matches the experiment design.
Why
Google recommends canonical links rather than `noindex` for alternate test URLs.

Do redirect-based temporary experiments use a temporary redirect rather than a permanent one?

Priority: 4
What
Confirmed platform requirement: Preserve the original URL as the intended long-term destination.
How
Use the documented temporary status and validate caching behavior.
Why
A permanent redirect sends the wrong persistence signal for a test.

Are minimum duration, collection volume, and maximum duration fixed before the experiment starts?

Priority: 5
What
Evidence-backed tactic: Avoid ending on a favorable fluctuation or leaving variants indefinitely.
How
Base duration on crawl/recrawl latency, variance, and business risk.
Why
Engines may take time to discover changes and measurements are stochastic.

Are campaigns, news, product launches, holidays, and demand shifts annotated and modeled?

Priority: 4
What
Evidence-backed tactic: Identify external events that affect both prompts and citations.
How
Use concurrent controls and event annotations.
Why
Changes in the world may affect visibility even when the page is unchanged.

For non-random tests, is treatment change compared with control change under a checked parallel-trends assumption?

Priority: 3
What
Experimental practice: Estimate relative change rather than treatment-only before/after change.
How
Plot pre-trends, disclose violations, and run sensitivity checks.
Why
Quasi-experiments can improve inference but rely on strong assumptions.

Are experiments reported as treatment effects with uncertainty rather than only percent lift?

Priority: 5
What
Evidence-backed tactic: Show absolute and relative change, interval, sample size, and baseline.
How
Use the precommitted estimator and retain null or negative results.
Why
Percent lift alone exaggerates small baselines and hides uncertainty.

Are many prompts, platforms, metrics, and segments accounted for when declaring a win?

Priority: 4
What
Evidence-backed tactic: Preselect primary outcomes and control false discoveries.
How
Use an appropriate correction or clearly label exploratory analyses.
Why
Testing enough slices will produce chance “wins.”

Are human graders blinded to treatment, date, and desired outcome when feasible?

Priority: 4
What
Evidence-backed tactic: Reduce expectancy bias in qualitative scoring.
How
Randomize answer order and remove treatment identifiers.
Why
Knowing which variant “should” win can influence ratings.

Are SEO traffic, conversions, accessibility, legal accuracy, and brand-safety guardrails monitored with rollback criteria?

Priority: 5
What
Evidence-backed tactic: Protect critical outcomes while testing GEO changes.
How
Set thresholds, owner, rollback mechanism, and evidence-preservation step before launch.
Why
A visibility experiment can harm users or established search performance.

Are hypotheses, variants, dates, units, metrics, results, and decisions stored in a searchable registry?

Priority: 4
What
Evidence-backed tactic: Preserve institutional learning and prevent repeated tests.
How
Require a registry ID in deployment notes and dashboards.
Why
Unrecorded tests create unexplained trend breaks and selective memory.

Are Google and Bing AI-report exports archived with property, filters, timezone, and export date?

Priority: 4
What
Evidence-backed tactic: Preserve first-party observations beyond changing dashboards.
How
Store immutable exports and a manifest without scraping unsupported fields.
Why
Limited retention, rollout, and product changes can break trend reconstruction.

Are site releases, content updates, crawler-policy changes, and known platform changes overlaid on trends?

Priority: 5
What
Evidence-backed tactic: Maintain a shared change ledger.
How
Link deployments and incidents to measurement windows.
Why
Unannotated changes encourage false causal stories.

Do visibility alerts account for expected variance, sample size, and repeated breaches?

Priority: 5
What
Evidence-backed tactic: Alert on material, sustained deviation rather than every point change.
How
Calibrate thresholds from baseline intervals and include data-quality gates.
Why
Deterministic thresholds create alert fatigue on stochastic systems.

Are GEO incidents classified by factual, legal, safety, identity, reach, and business impact with response targets?

Priority: 5
What
Evidence-backed tactic: Create severity levels, owners, and escalation routes.
How
Map critical claim types to acknowledgement, investigation, and remediation targets.
Why
A false address and a false medical or legal claim should not share one queue.

Is there a documented path to correct owned facts and report persistent answer-engine errors?

Priority: 5
What
Evidence-backed tactic: Separate source repair, indexing request, platform feedback, and stakeholder communication.
How
Preserve evidence, update the authoritative page, verify recrawl, and resample before closure.
Why
Repeated prompting without correcting source truth is not remediation.

Are prompt, response, citations, screenshots, source snapshots, timestamps, and context preserved for material incidents?

Priority: 5
What
Evidence-backed tactic: Create an auditable incident bundle.
How
Use controlled storage, legal retention rules, and immutable identifiers.
Why
Stochastic answers may disappear before investigation.

Are high-impact facts proactively tested with adversarial, ambiguous, and outdated formulations?

Priority: 5
What
Evidence-backed tactic: Probe identity, leadership, pricing, guarantees, safety, legal status, and crisis narratives.
How
Run a risk-ranked bilingual suite and route failures to qualified owners.
Why
Severe misinformation may be rare in a general benchmark.

Are superlatives, performance claims, comparisons, guarantees, and statistics backed by current evidence and conditions?

Priority: 5
What
Evidence-backed tactic: Maintain claim substantiation and expiry records.
How
Link approved claims to evidence and remove or update stale statements across languages.
Why
Answer engines can repeat unsupported marketing language as fact.

Are paid, employee, affiliate, gifted, or otherwise material relationships disclosed clearly near endorsements?

Priority: 5
What
Confirmed legal requirement: Keep endorsements honest and non-misleading, and disclose relationships that could materially affect how an audience evaluates them.
How
Make each disclosure clear and conspicuous in its actual format and context. Test placement, prominence, wording, and repetition rather than relying on a generic notice.
Why
The FTC treats an undisclosed material relationship as a deception risk and provides no universal wording or placement safe harbor.

Has counsel assessed whether AI interactions, deepfakes, or public-interest text fall within EU AI Act Article 50 duties?

Priority: 5
What
Confirmed legal requirement: Treat 2 August 2026 as the general application date after a role, content, audience, jurisdiction, and exemption analysis.
How
Record provider/deployer status, human-review process, disclosure method, and system placement date. The grace to 2 December 2026 is limited to systems placed on the market before 2 August and the Article 50(2) machine-readable marking and detection duty.
Why
The duties are material and date-sensitive but contextual. They do not require a blanket label for every AI-assisted edit.

Where an EU public-interest text exemption relies on human review or editorial control, is that review substantive and documented?

Priority: 5
What
Confirmed legal requirement: Preserve reviewer expertise, factual checks, approval authority, and editorial responsibility.
How
Require content-level approval records rather than only spelling or grammar checks.
Why
The Commission says superficial procedural checks are not sufficient human review.

Do high-risk synthetic or materially edited assets preserve verifiable provenance through publishing transformations?

Priority: 4
What
Evidence-backed tactic: Carry creation, edit, ingredient, and signer context where supported.
How
Validate the C2PA manifest after optimization, CDN processing, and download/re-upload flows.
Why
Provenance can help users and systems understand how an asset changed.

Does policy treat Content Credentials as provenance evidence rather than proof that an asset or claim is true?

Priority: 5
What
Evidence-backed tactic: Use a validated credential to assess signed provenance and asset integrity. A credential is not a truth score.
How
Review the signer, validation state, assertions, and content binding, then fact-check the depicted or stated claim independently.
Why
C2PA says valid manifests can coexist with misinformation. They establish verifiable association and freedom from tampering without establishing truth.

Are rights, licenses, attribution, quotation limits, and AI-use restrictions checked before publishing source-derived GEO content?

Priority: 5
What
Evidence-backed tactic: Maintain a rights record for text, images, datasets, testimonials, and generated assets.
How
Route uncertain reuse and training-origin questions to qualified counsel.
Why
Search visibility does not grant permission to copy or republish.

Are GEO vendors assessed for methodology, data rights, retention, security, model changes, exports, and incident duties?

Priority: 4
What
Evidence-backed tactic: Avoid outsourcing accountability to an opaque visibility score.
How
Contract for definitions, raw evidence access, change notice, deletion, and exit portability.
Why
Vendor methodology changes can silently rewrite historical results.

Where machine-readable AI licensing is part of the rights strategy, is every covered asset associated with a valid RSL 1.0 license?

Priority: 1
What
Experimental practice: Publish explicit permissions, prohibitions, licensing, attribution, or payment terms for automated uses.
How
Create an `application/rsl+xml` document in the RSL 1.0 namespace, scope its `<content>` rules carefully, and expose it through a supported discovery mechanism.
Why
RSL 1.0 defines an industry-recommendation format that distinguishes search, AI indexing and input, training, attribution, and payment terms.

Are RSL 1.0 associations validated for media type, scope, precedence, and cross-channel consistency?

Priority: 2
What
Experimental practice: Make the effective license unambiguous for every covered asset.
How
Test `application/rsl+xml`, absolute references, user-agent and asset scope, specific-over-broad precedence, and conflicts across robots.txt, HTTP `Link`, HTML, RSS, and embedded metadata.
Why
RSL requires clients to evaluate discoverable associations and applies specificity and restrictive conflict-resolution rules.

If Content Signals are used, do `search`, `ai-input`, and `ai-train` reflect separately approved choices rather than inherited assumptions?

Priority: 1
What
Experimental practice: Express post-access use preferences separately from crawler access rules.
How
Publish the Content Signals Policy and only approved `Content-Signal` values in robots.txt. Leave an undecided use unspecified and test the effective file after CDN composition.
Why
Cloudflare defines separate signals for search indexing, real-time AI input, and training. If a signal is absent, this mechanism neither grants nor restricts anything.

Where counsel selects TDMRep for applicable rights reservations, are the effective `tdm-reservation` and optional `tdm-policy` values valid and consistent?

Priority: 2
What
Experimental practice: Expose a machine-readable reservation of text-and-data-mining rights and, when offered, a route to licensing terms.
How
Publish `/.well-known/tdmrep.json` or supported HTTP, HTML, or asset metadata. Use `tdm-reservation: 1` to reserve rights, link a policy where applicable, and test precedence.
Why
The TDMRep Final Community Group Report defines interoperable reservation and policy discovery for lawfully accessible web content.

Is every property verified in Google Search Console?

Priority: 5
What
Confirmed platform requirement: Establish official crawl/index diagnostics for Google AI eligibility.
How
verify all protocols/subdomains or a suitable domain property and assign least-privilege access.
Why
Google directs site owners to Search Console for AI-feature technical diagnosis.

Does Google URL Inspection show the intended fetched HTML, canonical, index state, and preview controls?

Priority: 5
What
Confirmed platform requirement: Observe what Google received.
How
inspect representative pages after every template or edge-policy release.
Why
Google recommends it when AI preview controls appear ineffective.

Are Google generative-AI impressions reconciled with overall Web performance and analytics conversions?

Priority: 4
What
Confirmed platform requirement: Avoid double-counting or mistaking a rollout gap for zero visibility.
How
document report availability, compare URL/country/time trends, and join landing outcomes carefully.
Why
Dedicated data remains included in overall performance.

Are Bing Site Explorer error, warning, excluded, noindex, robots, redirect, and malware cohorts monitored?

Priority: 4
What
Confirmed platform requirement: Detect sitewide retrieval regressions.
How
baseline counts by folder and alert on discontinuities.
Why
Bing exposes these crawl/index cohorts directly.

Are crawler logs retained with verified vendor identity, status, bytes, latency, cache result, and route?

Priority: 5
What
Evidence-backed tactic: Build ground truth for access incidents.
How
enrich edge/origin logs after IP/reverse-DNS verification and chart by product token.
Why
UA strings alone are spoofable and platform consoles are incomplete.

Does a scheduled synthetic matrix fetch key URLs as each supported crawler from relevant regions?

Priority: 5
What
Evidence-backed tactic: Catch WAF, CDN, locale, and response drift before visibility falls.
How
test robots plus representative HTML/assets using documented UA/IP-safe methods.
Why
Intended policy can diverge from effective transport.

Are crawler-specific `403`, `401`, `429`, `5xx`, challenge, and zero-byte spikes alerted quickly?

Priority: 5
What
Evidence-backed tactic: Detect access regressions by bot and layer.
How
set anomaly alerts and preserve sampled request/response traces.
Why
These failures block retrieval or throttle crawling.

Is every robots/WAF change revalidated after each platform's documented propagation interval?

Priority: 4
What
Evidence-backed tactic: Close the deployment loop.
How
capture before/after parser output, live requests, logs, and platform diagnostics at the right delay.
Why
OpenAI and Perplexity document up-to-about-24-hour adjustment periods. Google requires recrawl.

Is a stable, versioned prompt panel used to observe answer-surface eligibility and citation changes?

Priority: 3
What
Experimental practice: Measure real outputs across engines, locales, devices, and fresh/anonymous sessions.
How
freeze prompt intent, record date/model/surface/citations, and repeat without treating outcomes as rank truth.
Why
No publisher console covers all answer engines.

Does the stale-content incident runbook update/delete the source, transport the change, and verify removal on every relevant index?

Priority: 5
What
Evidence-backed tactic: Handle harmful outdated answers end to end.
How
correct or remove the page, return the right status/directive, update sitemap/IndexNow/feed, request recrawl, and use platform removal channels if necessary.
Why
No single robots change clears every cached/indexed surface.

Are platform documentation and crawler-token changes reviewed on a fixed cadence?

Priority: 4
What
Evidence-backed tactic: Detect renamed bots, new controls, changed IP endpoints, and reporting features.
How
diff official pages monthly and trigger policy review on semantic changes.
Why
Multiple reviewed pages changed in July 2026 alone.

Are Google property-level and page-filtered Generative AI impression totals interpreted using their different aggregation rules?

Priority: 4
What
Confirmed platform requirement: Preserve report scope and aggregation mode with exports.
How
Do not reconcile totals by simple summation across pages.
Why
Multiple results from one property can count differently after URL filtering.

Does the Google Generative AI report avoid attributing impressions to prompts or grounding queries it does not expose?

Priority: 5
What
Confirmed platform requirement: Use only documented dimensions: pages, countries, dates, and devices.
How
Label private prompt sampling as a separate dataset.
Why
Joining unrelated prompt tests to platform impressions creates false precision.

Are Bing grounding queries described as examples rather than a complete demand or keyword dataset?

Priority: 5
What
Confirmed platform requirement: Preserve Bing's sampling caveat in downstream reports.
How
Prohibit extrapolation to search volume or total prompt share.
Why
Sampled grounding queries do not define the full retrieval universe.

Is a Google Generative AI impression counted only according to the platform's displayed-link definition?

Priority: 5
What
Confirmed platform requirement: Keep first-party report semantics distinct from private answer mentions.
How
Copy the current definition into the metric dictionary and date it.
Why
Private “visibility” observations are not interchangeable with Search Console impressions.

Is Google Generative AI report availability recorded rather than treating an absent report as zero visibility?

Priority: 4
What
Confirmed platform requirement: Distinguish unavailable, insufficient-data, and zero-impression states.
How
Add an availability flag and screenshot the report status.
Why
The report is not available to every property.

Are the newest Google Generative AI data points marked provisional until stabilized?

Priority: 3
What
Confirmed platform requirement: Prevent premature incident or success conclusions.
How
Reconcile recent periods after the documented processing window.
Why
Google says newest data can be preliminary and later change.
Every item shows its source, and the page carries the date it was last verified. That date does more work here than it would elsewhere. Crawler names, access controls and reporting all get revised without much warning, so a check that was accurate then can quietly stop being accurate now. None of it guarantees a citation. If classic search is the bigger gap right now, the SEO checklist is the other half of this.
Tell us what you're fixing