GEO for Shopify.
Readable, attributable, current.
Whether generative engines will cite this store when composing a recommendation. Generative engines compose an answer from sources they were allowed to read. GEO is whether the store is one of them, and whether a claim in the answer can be traced back to it.
- Target
- TGT-03 · GEO · Generative engine optimisation
- Checks
- 5 in pillar layer v1.0
- Weight sum
- 82 — relative within this pillar; not rescaled to 100
- Heaviest
- llms.txt presence and quality (20)
- Fix classes
- 0 automatic · 4 assisted · 1 guided — the design
- Crawler rule
- 6 points per fully disallowed AI agent, capped at 30, counted over all 10 agents the scan checks. The same penalty is subtracted from the composite.
- Composite
- A separate layer (rubric v1.0). Pillars do not sum to it.
- Fixes
- Planned automatic · Planned assisted
- Verify · rollback
- Planned · Planned
Nothing to draw on.
A crawler that robots.txt disallows does not read the store. What that costs depends on what the crawler feeds — a retrieval crawler’s search answers, a training crawler’s future models — but either way that crawler carries nothing of the store back to the system it serves. The score treats every disallowed AI agent the same way: a capped penalty, not an exclusion. Each subtracts 6 points from this pillar and from the composite, to a maximum of 30. Where a product names no vendor and carries no date, a model has less to attribute a claim to than it has with a record that does.
One file for access.
Four fields for attribution.
Where GEO lives on a Shopify store.
A generative engine writes an answer rather than listing links, from material it was allowed to read — at training time, or when the question is asked. For a Shopify store GEO comes down to three questions: may the engine’s crawlers read the store, can a claim in the catalogue be attributed to it, and does the record look maintained.
The first is decided by one file. A Shopify store’s /robots.txt applies to every crawler that visits; Shopify generates it by default and it is customised through robots.txt.liquid — so one pasted block of AI-crawler rules applies to the whole catalogue at once. The other two are decided by fields the merchant fills or leaves blank: vendor, the units in a description, the materials and standards it names, and updated_at.
The term comes from research — “GEO: Generative Engine Optimization” (Aggarwal et al., accepted to KDD 2024). Rank Sniper uses its own definition, above, and scores only what a scan of the storefront can observe.
Three gates
between a crawler and your catalogue.
Policy, delivery and use. Rank Sniper reads the first; the other two are stated so the score is not read as more than it is.
- GATE-01Available
Policy — robots.txt
What the store asks each crawler to do, per user-agent group. Rank Sniper reads it for 10 named AI agents and scores this pillar from it. It is a request, not a lock: RFC 9309 says the Robots Exclusion Protocol is not a substitute for content security, and OpenAI, Perplexity and Google document fetchers acting on a user’s request that may not follow robots.txt.
- GATE-02Planned
Delivery — what the server returns
A firewall, bot-management rule or CDN in front of a storefront can refuse a crawler that robots.txt allows. The public scan reads policy; it does not fetch the store as each AI agent to test delivery. Re-fetching as each declared agent belongs to the verification capability.
- GATE-03Vendor-documented · not observable
Use — what the operator does with it
Training, a live search index, a one-off fetch for a user, or a control token that governs use without crawling at all. The category decides what a block costs in the world, and it is set by each vendor’s documentation — not by anything a scan can see.
10 agents checked.
Six kinds of reader.
The 10 AI agents the scan checks, grouped by what their operators document them for. The penalty counts all of them the same way.
| Search | AI training | AI retrieval and search index | Assistant / user-initiated fetch | Control token (not a crawler) | Other | |
|---|---|---|---|---|---|---|
| Checked by the scan | Bytespider | GPTBot, ClaudeBot | OAI-SearchBot, PerplexityBot | ChatGPT-User | Google-Extended, Applebot-Extended | anthropic-ai, CCBot |
| Documented purpose | Indexing for a search engine. ByteDance documents Bytespider as the Toutiao Search crawler. | Collecting content that may be used to train foundation models (OpenAI, Anthropic). | Surfacing and linking sites in search answers (OpenAI, Perplexity). | Fetching a page because a user asked (OpenAI). | Not a crawler: a token governing whether content another crawler fetched may be used for training — and, for Google, grounding. | Common Crawl’s open crawl repository. anthropic-ai is absent from Anthropic’s current documentation. |
| Blocking it means | Pages leave that engine’s crawl. ByteDance’s page states no robots.txt policy. | Future content kept out of training data. Per OpenAI, not an opt-out from ChatGPT search. | The store is not shown in that engine’s search answers. OpenAI: it may still appear as a navigational link. | OpenAI states robots.txt rules may not apply to user-initiated fetches. | Google: no effect on inclusion in Google Search. Apple: Applebot-Extended does not crawl. | Depends on the operator. |
| Cost in the score | −6 each, capped at 30 | −6 each, capped at 30 | −6 each, capped at 30 | −6 each, capped at 30 | −6 each, capped at 30 | −6 each, capped at 30 |
| Crawler index | Search → | AI training → | AI retrieval and search index → | Assistant / user-initiated fetch → | Control token (not a crawler) → | Other → |
The GEO definition names four of them — GPTBot, ClaudeBot, PerplexityBot, Google-Extended — as its examples: two training crawlers, a retrieval crawler and a control token. The arithmetic does not use that list; the penalty counts every one of the 10 agents the scan checks. Agents outside that list — Claude-SearchBot, Claude-User and Perplexity-User among them — are not checked, and a store’s rules for them do not move the score.
A capped penalty,
not an exclusion.
What a disallowed AI agent costs in each score layer, computed from the constants the scorer uses.
- Per agent
- 6 points off the GEO pillar, and the same 6 off the composite.
- Cap
- 30 points, reached at 5 disallowed agents. Blocking more costs nothing further.
- Every agent blocked
- A catalogue that passes every composite check but disallows all 10 agents scores 66 on the composite instead of 96. The GEO pillar tops out at 70.
- By category
- No difference. A training crawler and a retrieval crawler cost the same. The penalty is a reading of policy, not a measured loss of visibility.
- Search crawlers
- Googlebot and Bingbot are not counted here. They penalise the SEO pillar only.
Blocking training crawlers is a legitimate decision. The score does not argue with it: it records that fewer AI systems may read the store, and leaves the trade-off with the merchant. The crawler penalty entry in the lexicon has the arithmetic for both layers.
5 checks.
Plus the crawler rule.
The GEO pillar as the scorer runs it today, read from the live derivation.
| Weight | Shareif all scored | Fix classdesign | What the scan reads | |
|---|---|---|---|---|
| llms.txt presence and quality | 20 | 24% | Guided | GET /llms.txt — a real text file, and whether it points at the catalogue |
| Vendor and provenance attribution | 16 | 20% | Assisted | A non-blank vendor on each product |
| Quotable claim density | 16 | 20% | Assisted | Numbers with units in the description (gsm, oz, cm) |
| Freshness (updated_at) | 16 | 20% | Assisted | updated_at within the last year |
| Named entities | 14 | 17% | Assisted | Named materials, standards or provenance in title or description |
| Sum of weights | 82 | — | — | Relative within this pillar; not rescaled to 100. |
The catalogue checks score the share of products that pass. llms.txt is one per store: a file that points at the catalogue earns the full check, a file that exists without doing so earns part of it, and a storefront page returned at /llms.txt counts as no file.
Share assumes every check was scored. A check whose input could not be read is unscored and leaves the denominator instead of counting against the store. The same rows, and how the pillar layer sits beside the composite, are on the rubric. The fix class column is the design — who will have to act — not a shipped fix.
llms.txt:
a proposal, scored as one.
The heaviest check in the pillar reads a file no major AI vendor documents consuming. Why it is scored, and exactly what that does not mean.
llms.txt is a proposal, published by Jeremy Howard in September 2024 and revised as a second version in 2026: a Markdown file at the domain root with a title, a short summary and lists of links to the pages a language model should read. It describes itself as a proposal to standardise — it is not a standard.
Google states that Google Search, including its generative features, ignores llms.txt, and that creating one will neither help nor harm a site’s visibility there. None of the crawler documentation from OpenAI, Anthropic or Perplexity cited on this page mentions consuming it.
Rank Sniper scores it at weight 20 for a narrow reason: it is the one place a store can state, in its own words, what it sells and where the catalogue is, instead of leaving every system to infer it. The check measures whether the store has done that. It is not evidence that any engine reads the file.
# Example Store > Linen bedding and throws, woven in Portugal. Prices in EUR. ## Catalogue - [All products](https://example.com/collections/all): every product, with sizes and materials - [Sitemap](https://example.com/sitemap.xml)
Publishing it is a guided fix: a file at the domain root sits outside the Shopify app’s write scope.
Attribution material.
What a scan can see.
Policy from robots.txt, terms from llms.txt, attribution from the catalogue record.
- robots.txt
- Groups for the 10 checked AI agents, parsed per group with path matching. A full disallow counts as a block.
- /llms.txt
- Presence; whether it is a real text file rather than a storefront page; whether it points at products, collections or a sitemap.
- Vendor
- The vendor field on each product — the brand a claim can be attributed to.
- Units
- Numbers with a unit in the description — gsm, oz, cm, mAh, % — facts specific enough to be repeated exactly.
- Named entities
- Materials, standards and provenance named in the title or description — merino, stainless, OEKO-TEX, “made in …”.
- updated_at
- The last-update timestamp in the public product listing, tested for a change within the last year.
Freshness reads a timestamp. It cannot tell a real change from a cosmetic one, so it rewards maintenance only as far as the record is honestly maintained. Named entities and units are what give a claim something to be attributed with; none of these checks can promise a citation.
Four robots.txt
and catalogue cases.
Illustrative, with the arithmetic computed from the scorer’s constants.
The blanket AI block
A robots.txt.liquid edit disallows GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot. That is 5 agents at 6 points: 30 off this pillar and off the composite, after the cap of 30. Googlebot is untouched, so SEO does not move.
Training out, search in
GPTBot and ClaudeBot are disallowed; OAI-SearchBot and PerplexityBot stay open. The score takes 12 points. Per OpenAI, ChatGPT search is governed by OAI-SearchBot, not GPTBot — the penalty records a policy choice, not a measured loss of search visibility.
A storefront page at /llms.txt
The store answers /llms.txt with its own HTML and a 200 status. The scan recognises an HTML document and counts no file, so llms.txt presence and quality scores zero.
The catalogue nobody touched
Products last edited two years ago fail Freshness (updated_at), and the result names examples. A real update — a corrected dimension, a new care note — is what the check is meant to track.
How GEO fails
before a word is written.
Disallowed
Policy asks the agent not to read the store. Scored: a capped penalty.
Refused at the edge
Policy allows; a firewall or CDN refuses. Not tested by the scan today.
Unattributable
No vendor, no units, no named entities — nothing to trace a claim to.
Stale
No change to the record in over a year.
Terms unstated
No llms.txt, so every system infers what the store sells and where.
Undocumented readers
Agents outside the checked list, whose rules the score does not read.
What happens today.
What is still planned.
Rank Sniper reads policy and scores attribution. Today the merchant changes robots.txt or the record and re-runs the scan.
- NOW-01Available
The scan reads and scores
Reads robots.txt first — for its own access and for the 10 AI agents — then the public product listing and /llms.txt, and scores the 5 GEO checks above.
- NOW-02Beta
The full report lists what fails
Every failing GEO check, heaviest-weighted first, in the full report sent by emailed sign-in link.
- NOW-03
You make the change, then re-scan
Publishing llms.txt states your terms and points at your catalogue, for any system that reads it. Naming a vendor and keeping updated_at current give a model something to attribute a claim to, and evidence that the record is maintained. Unblocking an agent in robots.txt lets that agent read the store again, and removes its 6-point share of the penalty.
- FIX-BPlanned
Assisted · 4 of 5 checks
Changes a native Shopify field — a title, description, tag, image or variant field. Drafted and staged; written only when you approve and apply it.
- FIX-C
Guided · 1 of 5 checks
Outside the app’s write scope — a theme template, a file at your domain root. You make the change from exact instructions; the next scan checks it. Until the app ships, the failing check and its fix class say where the change lives.
- VER-01Planned
Verification after a fix
Designed to re-fetch what agents actually receive once a change ships, and confirm it landed.
- VER-02Planned
Rollback
Designed to store the prior value of every write, with per-product restore.
GEO in
five lines.
What it is
Keeping a store readable by the systems that compose answers, and giving them attributable, current claims to compose from.
Why it matters
A generative engine can only draw on what its crawlers were allowed to read, and only attribute what names a source.
What Rank Sniper observes
5 checks — llms.txt, vendor, units, named entities, updated_at — plus robots.txt policy for 10 AI agents.
What Rank Sniper does today
Available Reads crawler policy and applies the capped penalty on every public scan. It does not test delivery, query an engine or measure citations.
What remains planned
Planned Re-fetching the store as each declared agent to see what it actually receives. Planned Drafted vendor, claim and entity changes for approval.
Access and attribution.
The rest depends on it.
SEO
Googlebot is SEO’s crawler — and, per Google, the robots.txt control for AI Overviews and AI Mode. Google-Extended, which GEO counts, does not affect Search. Google’s generative features in Search sit on the SEO side of that line.
Engage →AEO
GEO gets the store read and makes its claims attributable; AEO makes a passage quotable once it is read. Units are scored here, the opening sentence there.
Engage →AIO
GEO asks whether a claim has a brand behind it; AIO asks whether that brand is spelled consistently enough to be one entity. Both read the vendor field.
Engage →Scan a store.
See who may read it.
The public scan reads robots.txt for every checked AI agent and scores all four pillars.
- OfficialOverview of OpenAI crawlersOpenAI
- OfficialDoes Anthropic crawl data from the web, and how can site owners block the crawler?Anthropic
- OfficialPerplexity crawlersPerplexity
- OfficialGoogle’s common crawlers (incl. Google-Extended)Google Search Central
- OfficialGoogle’s user-triggered fetchersGoogle Search Central
- OfficialOptimizing your website for generative AI features on Google SearchGoogle Search Central
- OfficialAbout ApplebotApple
- OfficialCCBotCommon Crawl
- OfficialRFC 9309 — Robots Exclusion ProtocolIETF
- Officialrobots.txt.liquid templateShopify
- ExperimentalThe /llms.txt file — a proposalllmstxt.org
- ExperimentalGEO: Generative Engine Optimization (KDD 2024)arXiv
