TGT-03 · Pillar 3 of 4

GEO for Shopify.
Readable, attributable, current.

Whether generative engines will cite this store when composing a recommendation. Generative engines compose an answer from sources they were allowed to read. GEO is whether the store is one of them, and whether a claim in the answer can be traced back to it.

Intel // GEO readoutAvailable
Target
TGT-03 · GEO · Generative engine optimisation
Checks
5 in pillar layer v1.0
Weight sum
82 — relative within this pillar; not rescaled to 100
Heaviest
llms.txt presence and quality (20)
Fix classes
0 automatic · 4 assisted · 1 guided — the design
Crawler rule
6 points per fully disallowed AI agent, capped at 30, counted over all 10 agents the scan checks. The same penalty is subtracted from the composite.
Composite
A separate layer (rubric v1.0). Pillars do not sum to it.
Fixes
Planned automatic · Planned assisted
Verify · rollback
Planned · Planned
Stake · GEOWhat failing costs

Nothing to draw on.

A crawler that robots.txt disallows does not read the store. What that costs depends on what the crawler feeds — a retrieval crawler’s search answers, a training crawler’s future models — but either way that crawler carries nothing of the store back to the system it serves. The score treats every disallowed AI agent the same way: a capped penalty, not an exclusion. Each subtracts 6 points from this pillar and from the composite, to a maximum of 30. Where a product names no vendor and carries no date, a model has less to attribute a claim to than it has with a record that does.

01 // Shopify relevance

One file for access.
Four fields for attribution.

Where GEO lives on a Shopify store.

A generative engine writes an answer rather than listing links, from material it was allowed to read — at training time, or when the question is asked. For a Shopify store GEO comes down to three questions: may the engine’s crawlers read the store, can a claim in the catalogue be attributed to it, and does the record look maintained.

The first is decided by one file. A Shopify store’s /robots.txt applies to every crawler that visits; Shopify generates it by default and it is customised through robots.txt.liquid — so one pasted block of AI-crawler rules applies to the whole catalogue at once. The other two are decided by fields the merchant fills or leaves blank: vendor, the units in a description, the materials and standards it names, and updated_at.

The term comes from research — “GEO: Generative Engine Optimization” (Aggarwal et al., accepted to KDD 2024). Rank Sniper uses its own definition, above, and scores only what a scan of the storefront can observe.

02 // Access

Three gates
between a crawler and your catalogue.

Policy, delivery and use. Rank Sniper reads the first; the other two are stated so the score is not read as more than it is.

  1. GATE-01Available

    Policy — robots.txt

    What the store asks each crawler to do, per user-agent group. Rank Sniper reads it for 10 named AI agents and scores this pillar from it. It is a request, not a lock: RFC 9309 says the Robots Exclusion Protocol is not a substitute for content security, and OpenAI, Perplexity and Google document fetchers acting on a user’s request that may not follow robots.txt.

  2. GATE-02Planned

    Delivery — what the server returns

    A firewall, bot-management rule or CDN in front of a storefront can refuse a crawler that robots.txt allows. The public scan reads policy; it does not fetch the store as each AI agent to test delivery. Re-fetching as each declared agent belongs to the verification capability.

  3. GATE-03Vendor-documented · not observable

    Use — what the operator does with it

    Training, a live search index, a one-off fetch for a user, or a control token that governs use without crawling at all. The category decides what a block costs in the world, and it is set by each vendor’s documentation — not by anything a scan can see.

03 // Crawler categories

10 agents checked.
Six kinds of reader.

The 10 AI agents the scan checks, grouped by what their operators document them for. The penalty counts all of them the same way.

Categories per each operator’s own documentation · retrieved 2026-09-11 · Bytespider re-read 2026-09-12
SearchAI trainingAI retrieval and search indexAssistant / user-initiated fetchControl token (not a crawler)Other
Checked by the scanBytespiderGPTBot, ClaudeBotOAI-SearchBot, PerplexityBotChatGPT-UserGoogle-Extended, Applebot-Extendedanthropic-ai, CCBot
Documented purposeIndexing for a search engine. ByteDance documents Bytespider as the Toutiao Search crawler.Collecting content that may be used to train foundation models (OpenAI, Anthropic).Surfacing and linking sites in search answers (OpenAI, Perplexity).Fetching a page because a user asked (OpenAI).Not a crawler: a token governing whether content another crawler fetched may be used for training — and, for Google, grounding.Common Crawl’s open crawl repository. anthropic-ai is absent from Anthropic’s current documentation.
Blocking it meansPages leave that engine’s crawl. ByteDance’s page states no robots.txt policy.Future content kept out of training data. Per OpenAI, not an opt-out from ChatGPT search.The store is not shown in that engine’s search answers. OpenAI: it may still appear as a navigational link.OpenAI states robots.txt rules may not apply to user-initiated fetches.Google: no effect on inclusion in Google Search. Apple: Applebot-Extended does not crawl.Depends on the operator.
Cost in the score−6 each, capped at 30−6 each, capped at 30−6 each, capped at 30−6 each, capped at 30−6 each, capped at 30−6 each, capped at 30
Crawler indexSearch →AI training →AI retrieval and search index →Assistant / user-initiated fetch →Control token (not a crawler) →Other →

The GEO definition names four of them — GPTBot, ClaudeBot, PerplexityBot, Google-Extended — as its examples: two training crawlers, a retrieval crawler and a control token. The arithmetic does not use that list; the penalty counts every one of the 10 agents the scan checks. Agents outside that list — Claude-SearchBot, Claude-User and Perplexity-User among them — are not checked, and a store’s rules for them do not move the score.

04 // Crawler penalty

A capped penalty,
not an exclusion.

What a disallowed AI agent costs in each score layer, computed from the constants the scorer uses.

Per agent
6 points off the GEO pillar, and the same 6 off the composite.
Cap
30 points, reached at 5 disallowed agents. Blocking more costs nothing further.
Every agent blocked
A catalogue that passes every composite check but disallows all 10 agents scores 66 on the composite instead of 96. The GEO pillar tops out at 70.
By category
No difference. A training crawler and a retrieval crawler cost the same. The penalty is a reading of policy, not a measured loss of visibility.
Search crawlers
Googlebot and Bingbot are not counted here. They penalise the SEO pillar only.

Blocking training crawlers is a legitimate decision. The score does not argue with it: it records that fewer AI systems may read the store, and leaves the trade-off with the merchant. The crawler penalty entry in the lexicon has the arithmetic for both layers.

05 // What Rank Sniper evaluates

5 checks.
Plus the crawler rule.

The GEO pillar as the scorer runs it today, read from the live derivation.

GEO · pillar layer v1.0 · weights relative within this pillar; not rescaled to 100
WeightShareif all scoredFix classdesignWhat the scan reads
llms.txt presence and quality2024%GuidedGET /llms.txt — a real text file, and whether it points at the catalogue
Vendor and provenance attribution1620%AssistedA non-blank vendor on each product
Quotable claim density1620%AssistedNumbers with units in the description (gsm, oz, cm)
Freshness (updated_at)1620%Assistedupdated_at within the last year
Named entities1417%AssistedNamed materials, standards or provenance in title or description
Sum of weights82——Relative within this pillar; not rescaled to 100.

The catalogue checks score the share of products that pass. llms.txt is one per store: a file that points at the catalogue earns the full check, a file that exists without doing so earns part of it, and a storefront page returned at /llms.txt counts as no file.

Share assumes every check was scored. A check whose input could not be read is unscored and leaves the denominator instead of counting against the store. The same rows, and how the pillar layer sits beside the composite, are on the rubric. The fix class column is the design — who will have to act — not a shipped fix.

06 // llms.txt

llms.txt:
a proposal, scored as one.

The heaviest check in the pillar reads a file no major AI vendor documents consuming. Why it is scored, and exactly what that does not mean.

llms.txt is a proposal, published by Jeremy Howard in September 2024 and revised as a second version in 2026: a Markdown file at the domain root with a title, a short summary and lists of links to the pages a language model should read. It describes itself as a proposal to standardise — it is not a standard.

Google states that Google Search, including its generative features, ignores llms.txt, and that creating one will neither help nor harm a site’s visibility there. None of the crawler documentation from OpenAI, Anthropic or Perplexity cited on this page mentions consuming it.

Rank Sniper scores it at weight 20 for a narrow reason: it is the one place a store can state, in its own words, what it sells and where the catalogue is, instead of leaving every system to infer it. The check measures whether the store has done that. It is not evidence that any engine reads the file.

Illustrative · a minimal llms.txt
# Example Store
> Linen bedding and throws, woven in Portugal. Prices in EUR.

## Catalogue
- [All products](https://example.com/collections/all): every product, with sizes and materials
- [Sitemap](https://example.com/sitemap.xml)

Publishing it is a guided fix: a file at the domain root sits outside the Shopify app’s write scope.

07 // Observable signals

Attribution material.
What a scan can see.

Policy from robots.txt, terms from llms.txt, attribution from the catalogue record.

robots.txt
Groups for the 10 checked AI agents, parsed per group with path matching. A full disallow counts as a block.
/llms.txt
Presence; whether it is a real text file rather than a storefront page; whether it points at products, collections or a sitemap.
Vendor
The vendor field on each product — the brand a claim can be attributed to.
Units
Numbers with a unit in the description — gsm, oz, cm, mAh, % — facts specific enough to be repeated exactly.
Named entities
Materials, standards and provenance named in the title or description — merino, stainless, OEKO-TEX, “made in …”.
updated_at
The last-update timestamp in the public product listing, tested for a change within the last year.

Freshness reads a timestamp. It cannot tell a real change from a cosmetic one, so it rewards maintenance only as far as the record is honestly maintained. Named entities and units are what give a claim something to be attributed with; none of these checks can promise a citation.

08 // Shopify examples

Four robots.txt
and catalogue cases.

Illustrative, with the arithmetic computed from the scorer’s constants.

EX-01 · Illustrative

The blanket AI block

A robots.txt.liquid edit disallows GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot. That is 5 agents at 6 points: 30 off this pillar and off the composite, after the cap of 30. Googlebot is untouched, so SEO does not move.

EX-02 · Illustrative

Training out, search in

GPTBot and ClaudeBot are disallowed; OAI-SearchBot and PerplexityBot stay open. The score takes 12 points. Per OpenAI, ChatGPT search is governed by OAI-SearchBot, not GPTBot — the penalty records a policy choice, not a measured loss of search visibility.

EX-03 · Illustrative

A storefront page at /llms.txt

The store answers /llms.txt with its own HTML and a 200 status. The scan recognises an HTML document and counts no file, so llms.txt presence and quality scores zero.

EX-04 · Illustrative

The catalogue nobody touched

Products last edited two years ago fail Freshness (updated_at), and the result names examples. A real update — a corrected dimension, a new care note — is what the check is meant to track.

09 // Failure modes

How GEO fails
before a word is written.

FM-01

Disallowed

Policy asks the agent not to read the store. Scored: a capped penalty.

FM-02

Refused at the edge

Policy allows; a firewall or CDN refuses. Not tested by the scan today.

FM-03

Unattributable

No vendor, no units, no named entities — nothing to trace a claim to.

FM-04

Stale

No change to the record in over a year.

FM-05

Terms unstated

No llms.txt, so every system infers what the store sells and where.

FM-06

Undocumented readers

Agents outside the checked list, whose rules the score does not read.

10 // Remediation and verification

What happens today.
What is still planned.

Rank Sniper reads policy and scores attribution. Today the merchant changes robots.txt or the record and re-runs the scan.

  1. NOW-01

    The scan reads and scores

    Reads robots.txt first — for its own access and for the 10 AI agents — then the public product listing and /llms.txt, and scores the 5 GEO checks above.

    Available
  2. NOW-02

    The full report lists what fails

    Every failing GEO check, heaviest-weighted first, in the full report sent by emailed sign-in link.

    Beta
  3. NOW-03

    You make the change, then re-scan

    Publishing llms.txt states your terms and points at your catalogue, for any system that reads it. Naming a vendor and keeping updated_at current give a model something to attribute a claim to, and evidence that the record is maintained. Unblocking an agent in robots.txt lets that agent read the store again, and removes its 6-point share of the penalty.

  4. FIX-B

    Assisted · 4 of 5 checks

    Changes a native Shopify field — a title, description, tag, image or variant field. Drafted and staged; written only when you approve and apply it.

    Planned
  5. FIX-C

    Guided · 1 of 5 checks

    Outside the app’s write scope — a theme template, a file at your domain root. You make the change from exact instructions; the next scan checks it. Until the app ships, the failing check and its fix class say where the change lives.

  6. VER-01

    Verification after a fix

    Designed to re-fetch what agents actually receive once a change ships, and confirm it landed.

    Planned
  7. VER-02

    Rollback

    Designed to store the prior value of every write, with per-product restore.

    Planned
11 // Field brief

GEO in
five lines.

FB-01

What it is

Keeping a store readable by the systems that compose answers, and giving them attributable, current claims to compose from.

FB-02

Why it matters

A generative engine can only draw on what its crawlers were allowed to read, and only attribute what names a source.

FB-03

What Rank Sniper observes

5 checks — llms.txt, vendor, units, named entities, updated_at — plus robots.txt policy for 10 AI agents.

FB-04

What Rank Sniper does today

Available Reads crawler policy and applies the capped penalty on every public scan. It does not test delivery, query an engine or measure citations.

FB-05

What remains planned

Planned Re-fetching the store as each declared agent to see what it actually receives. Planned Drafted vendor, claim and entity changes for approval.

13 // Engage

Scan a store.
See who may read it.

The public scan reads robots.txt for every checked AI agent and scores all four pillars.

Shopify GEO: generative engine optimisation for AI answers — Rank Sniper