Intel // Crawler index

33 agents indexed.
10 checked by name.

The crawler and control tokens that search, AI and shopping operators document: who runs each, what it is for, whether it obeys robots.txt and what blocking it changes, every fact linked to the operator’s own page. Every Rank Sniper scan checks 10 of them by name. A block is a capped penalty in the score, and each card says whether that agent counts.

Intel // Crawler index readoutAvailable
Entries indexed
33 user-agents and control tokens
Operators
13: Google, Microsoft, Apple, OpenAI, Anthropic, Meta, Amazon, Mistral AI, Perplexity, DuckDuckGo, Common Crawl, ByteDance, Rank Sniper
Checked by every scan
10: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, anthropic-ai, PerplexityBot, Google-Extended, Applebot-Extended, Bytespider, CCBot
Read for the SEO pillar
Googlebot, Bingbot
Evidence
32 verified on the operator’s own page · 1 secondary · 0 research needed
Penalty rule
−6 per blocked listed agent, capped at −30 (reached at 5 agents), on the composite and the GEO pillar · rubric v1.0
Last reviewed
2026-09-11
01 // Classes

Four decisions,
not one list.

An AI crawler is not one thing. The class of an agent decides what a block in robots.txt actually changes, so read the class before the brand.

KEY-01Read this first

Training is not retrieval.

A training crawler collects content that may train a model. An AI search crawler builds the index an answer is retrieved from. A user-initiated fetch reads one page because a person asked. A control token crawls nothing and only states how fetched content may be used. Blocking training while allowing search is a defensible choice; blocking everything by accident is not. Each user-agent below is filed under one class.

Crawler classes · counted from 33 entries · reviewed 2026-09-11
What it doesWhat a block changesper operatorrobots.txtper operatorRank Snipertoday
SearchCrawl pages for a search index.Pages leave that engine’s crawl. For Googlebot, Google says this is also the control for its AI features.3 honour · 1 not documented1 of 4 checked by every scan; 2 read for the SEO pillar
AI trainingCollect content that may train foundation models.Asks the operator to leave future content out of training. Their search agents are separate.5 honour2 of 5 checked by every scan
AI retrieval and search indexIndex pages for an AI search or answer product.Per OpenAI and Anthropic, less or no visibility in that product’s search answers.5 honour · 2 not documented2 of 7 checked by every scan
Assistant / user-initiated fetchFetch one page because a user asked an assistant to.Depends on the operator. Several say their fetcher may ignore the rule.1 honour · 5 may not follow · 1 not documented1 of 7 checked by every scan
Commerce / shoppingCrawl for shopping surfaces.Pages are withheld from that shopping crawl.1 honour0 of 1 checked by every scan
Control token (not a crawler)State a use decision about content another crawler fetched.Opts out of training (Google adds grounding). Both operators say search is unaffected.2 honour2 of 2 checked by every scan
OtherVaries: archive building, link previews, ad checks, generic fetching.Depends on the agent; each card says what its operator documents.5 honour · 1 may not follow · 1 not documented2 of 7 checked by every scan
03 // AI training

Training crawlers.
A consent decision.

Crawlers whose operators say the content may be used to train models. By the operators’ own accounts, blocking one is an opt-out from training, not from their search products, which run under other tokens.

OpenAIEvidence: verified

GPTBot

Crawls content that may be used to train OpenAI’s generative AI foundation models.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Published by the operator: JSON list.
Rank Sniper
Checked by every scan — counts toward the capped penalty.

Disallowing it tells OpenAI a site’s content should not be used in training. OpenAI says each of its crawler settings is independent, so this block does not remove a store from ChatGPT search.

Official · OpenAI
AnthropicEvidence: verified

ClaudeBot

Collects web content that could potentially contribute to Anthropic’s model training.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Published by the operator: JSON list.
Rank Sniper
Checked by every scan — counts toward the capped penalty.

Blocking it asks Anthropic to exclude the site’s future materials from training datasets. Claude’s search and user-directed fetches run under other agents.

Anthropic supports the non-standard Crawl-delay extension and says its bots will not attempt to bypass CAPTCHAs.

Official · Anthropic
MetaEvidence: verified

meta-externalagent

Crawls the web for use cases such as training foundation AI models or improving products by indexing content directly.

robots.txt
Honours robots.txt, per the operator.Meta: “In order to block these crawlers, add a disallow for the relevant crawler to robots.txt.” It is not among the agents Meta says may bypass robots.txt.
IP ranges
Not documented on the page read.
Rank Sniper
Not checked by the scan today.

The Meta agent whose stated purpose includes training foundation models.

Meta asks for up to 24 hours for robots.txt changes to take effect.

Official · Meta
AmazonEvidence: verified

Amazonbot

Used to improve Amazon’s products and services; its content may be used to train AI models.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Published by the operator: listed on its page.
Rank Sniper
Not checked by the scan today.

Amazon runs search under a separate token, Amzn-SearchBot, so a training decision here is not a search decision.

Amazon’s crawlers do not support Crawl-delay and may cache robots.txt for up to 30 days.

Official · Amazon
Mistral AIEvidence: verified

MistralAI-Training

Builds Mistral’s training datasets.

robots.txt
Honours robots.txt, per the operator.Mistral: “Webmasters can disallow this user agent in their robots.txt file.”
IP ranges
Not documented on the page read.
Rank Sniper
Not checked by the scan today.

Mistral splits training, search indexing and user fetches into three tokens, so each can be decided alone.

Official · Mistral AI
04 // AI retrieval and search index

AI search crawlers.
The index behind the answer.

Crawlers that build the index an AI search or answer product retrieves from. Every operator here documents them apart from training, so a store can allow one and refuse the other.

OpenAIEvidence: verified

OAI-SearchBot

Surfaces websites in ChatGPT search results.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Published by the operator: JSON list.
Rank Sniper
Checked by every scan — counts toward the capped penalty.

OpenAI says opted-out sites will not be shown in ChatGPT search answers, though they can still appear as navigational links.

OpenAI says a robots.txt change takes about 24 hours to take effect.

Official · OpenAI
AnthropicEvidence: verified

Claude-SearchBot

Navigates the web to improve search result quality for Claude’s users.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Published by the operator: JSON list.
Rank Sniper
Not checked by the scan today.

Anthropic says blocking it may reduce a site’s visibility in user search results.

Official · Anthropic
PerplexityEvidence: verified

PerplexityBot

Surfaces and links websites in Perplexity’s search results; Perplexity says it is not used to crawl content for AI foundation models.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Published by the operator: JSON list.
Rank Sniper
Checked by every scan — counts toward the capped penalty.

By Perplexity’s own description, blocking it is a search decision, not a training one.

Official · Perplexity
MetaEvidence: verified

meta-webindexer

Navigates the web to improve Meta AI search result quality.

robots.txt
Not documented by the operator.
IP ranges
Not documented on the page read.
Rank Sniper
Not checked by the scan today.

Meta separates this search indexer from meta-externalagent, the agent whose purpose includes training.

Official · Meta
AmazonEvidence: verified

Amzn-SearchBot

Crawls for Amazon search experiences such as Alexa; Amazon says it does not crawl content for generative AI model training.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Published by the operator: listed on its page.
Rank Sniper
Not checked by the scan today.

The Amazon agent to keep allowed if a store blocks Amazonbot for training but wants to stay in Amazon’s search experiences.

Official · Amazon
Mistral AIEvidence: verified

MistralAI-Index

Indexes pages for Mistral’s search; Mistral says it is “not used for generative AI training of any kind.”

robots.txt
Not documented by the operator.
IP ranges
Published by the operator: JSON list.
Rank Sniper
Not checked by the scan today.

Mistral’s search index, kept apart from its training crawler.

Mistral’s page does not state explicitly whether this agent obeys robots.txt.

Official · Mistral AI
DuckDuckGoEvidence: verified

DuckAssistBot

Crawls in real time for DuckDuckGo’s AI-assisted answers; DuckDuckGo says the data is not used in any way to train AI models.

User-Agent
DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)
robots.txt
Honours robots.txt, per the operator.DuckDuckGo: a robots.txt disallow takes effect within 72 hours.
IP ranges
Published by the operator: JSON list.
Rank Sniper
Not checked by the scan today.

DuckDuckGo says opting out does not affect organic search rankings.

Official · DuckDuckGo
05 // Assistant / user-initiated fetch

User-initiated fetchers.
robots.txt may not apply.

Fetches made because a person asked an assistant to read a page. OpenAI, Perplexity, Google, Meta and Amazon each say robots.txt may not apply, or not fully, to theirs. Anthropic says its bots honour it.

OpenAIEvidence: verified

ChatGPT-User

Fetches pages for user actions in ChatGPT and Custom GPTs, including GPT Actions; not used to crawl the web automatically.

robots.txt
May not follow robots.txt, per the operator.OpenAI: “Because these actions are initiated by a user, robots.txt rules may not apply.”
IP ranges
Published by the operator: JSON list.
Rank Sniper
Checked by every scan — counts toward the capped penalty.

OpenAI says it is not used to determine whether content may appear in ChatGPT search.

Official · OpenAI
AnthropicEvidence: verified

Claude-User

Accesses websites when people ask Claude questions.

robots.txt
Honours robots.txt, per the operator.Anthropic’s blanket statement that its bots honour robots.txt covers this agent. Unlike OpenAI, Perplexity and Google, Anthropic states no user-initiated exception.
IP ranges
Published by the operator: JSON list.
Rank Sniper
Not checked by the scan today.

Anthropic says blocking it may reduce a site’s visibility for user-directed web search.

Official · Anthropic
PerplexityEvidence: verified

Perplexity-User

Visits a page to help answer a user’s question, and may link to it in the response.

robots.txt
May not follow robots.txt, per the operator.Perplexity: “Since a user requested the fetch, this fetcher generally ignores robots.txt rules.”
IP ranges
Published by the operator: JSON list.
Rank Sniper
Not checked by the scan today.

A robots.txt rule is not a dependable way to stop it; the vendor says so.

Official · Perplexity
MetaEvidence: verified

meta-externalfetcher

Fetches individual links at a user’s request and supports product functions such as evaluating and improving agentic AI capabilities.

robots.txt
May not follow robots.txt, per the operator.Meta: it “may bypass robots.txt because it performs fetches that were requested by the user.”
IP ranges
Not documented on the page read.
Rank Sniper
Not checked by the scan today.

Meta’s user-requested fetcher, documented separately from its training and search agents.

Official · Meta
AmazonEvidence: verified

Amzn-User

Fetches live information for user actions, for example through Alexa.

robots.txt
May not follow robots.txt, per the operator.Amazon: “Because actions taken by Amzn-User can be initiated by a user, it may not follow all robots.txt directives.”
IP ranges
Published by the operator: listed on its page.
Rank Sniper
Not checked by the scan today.

Amazon’s live, user-driven fetcher; distinct from both Amazonbot and Amzn-SearchBot.

Official · Amazon
Mistral AIEvidence: verified

MistralAI-User

Fetches pages for user actions in Mistral’s Vibe; not automatic crawling and not training.

robots.txt
Not documented by the operator.Mistral: “MistralAI-User governs which sites these user requests can be made to.” Whether it obeys robots.txt is not stated explicitly.
IP ranges
Published by the operator: JSON list.
Rank Sniper
Not checked by the scan today.

Mistral’s user-driven fetcher; its robots.txt behaviour is the one unverified fact on this card.

Official · Mistral AI
GoogleEvidence: verified

Google-Agent

Navigates the web and performs user-initiated actions, with mobile and desktop variants.

robots.txt
May not follow robots.txt, per the operator.Google says of its user-triggered fetchers: “Because the fetch was requested by a user, these fetchers generally ignore robots.txt rules.”
IP ranges
Published by the operator: listed on its page.
Rank Sniper
Not checked by the scan today.

Google’s agent for acting on a user’s behalf, as opposed to crawling for the index.

Google lists other user-triggered fetchers under the same rule, among them Google-GeminiNotebook, Google-Read-Aloud and FeedFetcher-Google.

06 // Commerce / shopping

Shopping crawlers.
Where listings are read.

A crawler whose documented purpose is shopping surfaces. Of the operators read for this index, only Google documents one.

GoogleEvidence: verified

Storebot-Google

Crawls for all surfaces of Google Shopping, for example the Shopping tab in Google Search.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Published by the operator: listed on its page.
Rank Sniper
Not checked by the scan today.

The Google crawler most directly tied to product listings. A Shopify store that disallows it withholds its pages from that crawl.

07 // Control token (not a crawler)

Control tokens.
Not crawlers at all.

Names you write in robots.txt to decide how content another crawler fetched may be used. Google calls Google-Extended a standalone product token; Apple says Applebot-Extended does not crawl webpages.

GoogleEvidence: verified

Google-Extended

A standalone product token that manages whether content Google crawls may be used to train future Gemini models and for grounding. It is not a crawler.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Not applicable: a control token is not a crawler.
Rank Sniper
Checked by every scan — counts toward the capped penalty.

Google states it does not affect a site’s inclusion in Google Search and is not a ranking signal. AI Overviews and AI Mode are governed by Googlebot.

AppleEvidence: verified

Applebot-Extended

Lets a publisher opt out of its content being used to train Apple’s foundation models. Apple states it does not crawl webpages.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Not applicable: a control token is not a crawler.
Rank Sniper
Checked by every scan — counts toward the capped penalty.

Apple says pages disallowed for this token can still appear in its search results.

Official · Apple
08 // Other

Other agents.
Archives, previews, legacy tokens.

Agents that fit none of the classes above: an open crawl archive, ads and link-preview fetchers, generic Google crawlers, two tokens without current operator documentation, and Rank Sniper’s own scanner.

Common CrawlEvidence: verified

CCBot

Builds Common Crawl’s open repository of web crawl data.

User-Agent
CCBot/2.0 (https://commoncrawl.org/faq/)
robots.txt
Honours robots.txt, per the operator.Common Crawl: blocked with User-agent: CCBot and Disallow: /; it obeys Crawl-delay and honours nofollow.
IP ranges
Published by the operator: JSON list.
Rank Sniper
Checked by every scan — counts toward the capped penalty.

A block keeps a store out of an open crawl archive. Common Crawl does not describe CCBot as an AI-training crawler on the pages read, so this index does not either.

Common Crawl says it is aware of crawlers falsely identifying themselves as CCBot.

Official · Common Crawl
OpenAIEvidence: verified

OAI-AdsBot

Validates the safety of ad landing pages submitted to ChatGPT, visiting only submitted pages; not used to train foundation models.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Published by the operator: JSON list.
Rank Sniper
Not checked by the scan today.

It matters only to a store that advertises in ChatGPT, and visits nothing else.

Official · OpenAI
GoogleEvidence: verified

GoogleOther

A generic crawler that may be used by various Google product teams for fetching publicly accessible content.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Published by the operator: listed on its page.
Rank Sniper
Not checked by the scan today.

Google’s catch-all crawler; it is not Googlebot, so its rules do not decide Search.

GoogleOther-Image and GoogleOther-Video are separate variants.

GoogleEvidence: verified

Google-CloudVertexBot

Crawls requested by site owners for building Vertex AI Agents; Google says it has no effect on Google Search.

robots.txt
Honours robots.txt, per the operator.
IP ranges
Published by the operator: listed on its page.
Rank Sniper
Not checked by the scan today.

It visits because a site owner asked it to. A store that never set up a Vertex AI agent has little reason to meet it.

MetaEvidence: verified

facebookexternalhit

Fetches the title, description and thumbnail of content shared on Meta’s apps.

robots.txt
May not follow robots.txt, per the operator.Meta: it “might bypass robots.txt when performing security or integrity checks.”
IP ranges
Not documented on the page read.
Rank Sniper
Not checked by the scan today.

It builds the link preview when a product URL is shared on Meta’s apps: a social-sharing concern, not an AI one.

Official · Meta
AnthropicEvidence: secondary

anthropic-ai

An older Anthropic token that still appears in robots.txt files, which is why the scan checks it.

robots.txt
Not documented by the operator.
IP ranges
Not documented on the page read.
Rank Sniper
Checked by every scan — counts toward the capped penalty.

A rule written only for this token may match no agent Anthropic documents today. Name ClaudeBot, Claude-SearchBot and Claude-User explicitly.

Anthropic’s current crawler article does not mention it. Third parties describe it, and Claude-Web, as deprecated; that status is not documented by Anthropic.

Official · Anthropic
Rank SniperEvidence: verified

RankSniperBot

Fetches a handful of public storefront pages when someone requests a Rank Sniper scan of that store.

User-Agent
RankSniperBot/1.0 (+https://ranksniperhq.com/bot)
robots.txt
Honours robots.txt, per the operator.It reads robots.txt before anything else; a disallowed path is not fetched by any identity, and a refused catalogue ends the scan.
IP ranges
Not published.
Rank Sniper
This is the scanner itself; it is not scored.

The crawler behind the free scan. Two lines of robots.txt turn it off.

Where Shopify’s edge refuses a declared request, it retries that single request with a standard browser profile and records the fallback in the scan trace.

Verified · Rank Sniper
09 // Score

What a block costs.
A capped penalty.

How robots.txt feeds the score, read from the scoring code. The crawler penalty measures crawlability for the agents the scan names; it does not remove a store from the score.

Each agent on the scan’s list that your robots.txt fully disallows subtracts 6 points from the composite score after the catalogue and schema blend, capped at 30; the result is rounded, with a floor of 1. The same capped amount comes off the GEO pillar. Catalogue quality still counts: a store with one blocked agent loses 6 points, not its score.

Fully disallowed means the group that applies to the agent (its own group, or User-agent: * when no group names it) contains Disallow: / with no Allow: / to override it. A path rule such as Disallow: /collections/ does not count. A wildcard block counts against every listed agent the file does not name, so User-agent: * with Disallow: / reaches the cap on its own: 10 agents × 6 is 60, held to 30.

The penalty is the same for every listed agent, whatever its class. A deliberate opt-out through Google-Extended, a control token that Google says has no effect on Search, costs the same 6 points as an accidental block on a search crawler. Rubric v1.0 does not weigh an agent by what it does, so read the number with that in mind.

Googlebot and Bingbot blocks never touch the composite. They take the same 6-point, 30-cap penalty from the SEO pillar only (pillar layer v1.0). Pillar scores are a separate layer and do not sum to the composite; the rubric shows both, as the code computes them. For how the GEO and SEO pillars fit the four disciplines, see GEO and SEO in the lexicon.

Per blocked listed agent−6
Cap−30
Agents to reach the cap5
Listed AI agents10
SEO pillar onlyGooglebot · Bingbot
10 // Policy vs delivery

Policy is not delivery.
robots.txt is a request.

robots.txt states a policy. Whether an agent actually receives your pages is decided somewhere else, and the two can disagree in both directions.

A firewall, a CDN or bot-management rule, or an app that filters traffic by user-agent decides what a crawler is actually served. The standard is explicit that the file is not an access control: “The Robots Exclusion Protocol is not a substitute for valid content security measures” (RFC 9309). Google says robots.txt “is not a mechanism for keeping a web page out of Google”, and points to noindex or password protection instead.

So a file that allows GPTBot means nothing if an edge rule refuses it, and a file that disallows Perplexity-User may not stop it, because Perplexity says that fetcher generally ignores robots.txt. User-agent strings can also be forged: Common Crawl warns of crawlers posing as CCBot. That is why most operators in this index publish IP ranges, and why a rule that trusts the name alone can be fooled.

Rank Sniper reads policy only. Testing delivery, which means requesting your pages as each agent and recording what comes back, is not built. The nearest item on the roadmap is post-fix verification (Post-fix verification — what agents actually receive: Planned), designed to re-fetch a page as each declared AI user agent after a fix ships. One indirect signal exists today: when Shopify’s edge refuses RankSniperBot’s declared request and the scan falls back to a browser profile, the scan trace records it. That says something about your edge, and nothing certain about how it treats GPTBot or ClaudeBot. How RankSniperBot identifies itself.

11 // Shopify robots.txt

Deciding your robots.txt.
On Shopify, in Liquid.

Where the file lives on a Shopify store, how crawlers resolve it, and what each common decision means, including what it costs in Rank Sniper’s score today. See also robots.txt and robots.txt.liquid in the lexicon.

Shopify generates a default robots.txt that, in its words, “works for most stores”. To change it, add a robots.txt.liquid template to the theme’s templates/ directory. It renders /robots.txt and exposes the robots, group, rule, user_agent and sitemap Liquid objects. Shopify strongly recommends building on those objects rather than hard-coding the file, because its default rules are updated regularly. The Shopify pages read for this index do not list the default rules, so read your own store’s /robots.txt before you edit it.

Resolution follows RFC 9309. A crawler obeys the group that names it and falls back to User-agent: * only when none does. Within a group, “the most specific match found MUST be used”: the rule with the most octets. When an Allow and a Disallow match equally, the Allow rule should be used. Rank Sniper’s parser matches tokens case-insensitively, so bingbot and Bingbot name one group.

If these tokens are disallowed · score effect computed from rubric v1.0
Tokens to nameWhat the operators sayper operatorScore if disallowedtoday
Keep content out of model trainingGPTBot, ClaudeBot, meta-externalagent, Amazonbot, MistralAI-Training, Google-Extended, Applebot-ExtendedA training opt-out. OpenAI, Anthropic, Amazon and Mistral each run search under a different token; Google and Apple say search is unaffected by their control tokens.composite and GEO pillar −24 (4 of these on the scan’s list)
Stay in AI search and retrievalOAI-SearchBot, Claude-SearchBot, PerplexityBot, meta-webindexer, Amzn-SearchBot, MistralAI-Index, DuckAssistBotLeave these allowed. OpenAI says opted-out sites are not shown in ChatGPT search answers; Anthropic says blocking may reduce visibility.composite and GEO pillar −12 (2 of these on the scan’s list)
Refuse user-initiated fetchesChatGPT-User, Claude-User, Perplexity-User, meta-externalfetcher, Amzn-User, Google-AgentOnly partly possible. OpenAI, Perplexity, Google, Meta and Amazon say robots.txt may not apply to these; Anthropic says it honours robots.txt.composite and GEO pillar −6 (1 of these on the scan’s list)
Stay in Google Search, AI Overviews and AI ModeGooglebotLeave Googlebot allowed. Google names its robots.txt rules as the control for Search, its AI features included; Google-Extended does not affect Search.SEO pillar −6, composite unaffected
Stay in Google Shopping’s crawlStorebot-GoogleLeave it allowed. Google says it crawls for all surfaces of Google Shopping.No score effect: none of these is checked today.
robots.txt output — a training opt-out that leaves AI search alone
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

One group, four tokens, every search and retrieval agent left to the default group. In Shopify, keep the default groups rendered through the Liquid objects and add this as an extra group. Today Rank Sniper would subtract 24 points for it, because every token in it is on the scan’s list. The score does not yet tell a deliberate training opt-out from an accidental block; the choice can still be right. robots.txt governs access only; llms.txt is a separate proposal, covered in the llms.txt guide for Shopify.

12 // Coverage

Not checked yet.
20 of 33 entries.

The scan checks 10 AI tokens and reads 2 search crawlers for the SEO pillar. Everything below is indexed here but not observed by the scan.

Adding a token to the scan changes the composite for every store that blocks it, so the list moves only with a rubric version bump (today v1.0) and an entry in the changelog, never quietly. RankSniperBot is listed above for completeness: it is how the scan reads, not something it scores. Its own rules are on /bot.

Sources · read 2026-09-11
  1. OfficialGoogle common crawlersGoogle Search Central
  2. OfficialAI features and your websiteGoogle Search Central
  3. OfficialIntroducing AI Performance in Bing Webmaster Tools (public preview)Microsoft Bing
  4. OfficialAnnouncing new options for webmasters to control usage of their content in Bing ChatMicrosoft Bing
  5. OfficialAbout ApplebotApple
  6. OfficialOverview of OpenAI crawlersOpenAI
  7. OfficialDoes Anthropic crawl data from the web, and how can site owners block the crawler?Anthropic
  8. OfficialMeta web crawlersMeta
  9. OfficialAmazonbotAmazon
  10. OfficialMistral AI robotsMistral AI
  11. OfficialPerplexity crawlersPerplexity
  12. OfficialDuckAssistBotDuckDuckGo
  13. OfficialGoogle user-triggered fetchersGoogle Search Central
  14. OfficialCCBotCommon Crawl
  15. OfficialAbout Bytespider (Toutiao Search, in Chinese)ByteDance · Toutiao
  16. VerifiedRankSniperBot: crawler identity and how to block itRank Sniper
  17. OfficialRFC 9309: Robots Exclusion ProtocolIETF
  18. OfficialIntroduction to robots.txtGoogle Search Central
  19. OfficialCustomize robots.txtShopify
  20. Officialrobots.txt.liquid templateShopify
  21. OfficialVerifying Googlebot and other Google crawlersGoogle Search Central
Crawler Index: 33 AI and search user-agents for Shopify stores — Rank Sniper