Common CrawlEvidence: verified
CCBot
Builds Common Crawl’s open repository of web crawl data.
- User-Agent
CCBot/2.0 (https://commoncrawl.org/faq/)- robots.txt
- Honours robots.txt, per the operator.Common Crawl: blocked with User-agent: CCBot and Disallow: /; it obeys Crawl-delay and honours nofollow.
- IP ranges
- Published by the operator: JSON list.
- Rank Sniper
- Checked by every scan — counts toward the capped penalty.
A block keeps a store out of an open crawl archive. Common Crawl does not describe CCBot as an AI-training crawler on the pages read, so this index does not either.
Common Crawl says it is aware of crawlers falsely identifying themselves as CCBot.
OpenAIEvidence: verified
OAI-AdsBot
Validates the safety of ad landing pages submitted to ChatGPT, visiting only submitted pages; not used to train foundation models.
- robots.txt
- Honours robots.txt, per the operator.
- IP ranges
- Published by the operator: JSON list.
- Rank Sniper
- Not checked by the scan today.
It matters only to a store that advertises in ChatGPT, and visits nothing else.
GoogleEvidence: verified
GoogleOther
A generic crawler that may be used by various Google product teams for fetching publicly accessible content.
- robots.txt
- Honours robots.txt, per the operator.
- Rank Sniper
- Not checked by the scan today.
Google’s catch-all crawler; it is not Googlebot, so its rules do not decide Search.
GoogleOther-Image and GoogleOther-Video are separate variants.
GoogleEvidence: verified
Google-CloudVertexBot
Crawls requested by site owners for building Vertex AI Agents; Google says it has no effect on Google Search.
- robots.txt
- Honours robots.txt, per the operator.
- Rank Sniper
- Not checked by the scan today.
It visits because a site owner asked it to. A store that never set up a Vertex AI agent has little reason to meet it.
MetaEvidence: verified
facebookexternalhit
Fetches the title, description and thumbnail of content shared on Meta’s apps.
- robots.txt
- May not follow robots.txt, per the operator.Meta: it “might bypass robots.txt when performing security or integrity checks.”
- IP ranges
- Not documented on the page read.
- Rank Sniper
- Not checked by the scan today.
It builds the link preview when a product URL is shared on Meta’s apps: a social-sharing concern, not an AI one.
AnthropicEvidence: secondary
anthropic-ai
An older Anthropic token that still appears in robots.txt files, which is why the scan checks it.
- robots.txt
- Not documented by the operator.
- IP ranges
- Not documented on the page read.
- Rank Sniper
- Checked by every scan — counts toward the capped penalty.
A rule written only for this token may match no agent Anthropic documents today. Name ClaudeBot, Claude-SearchBot and Claude-User explicitly.
Anthropic’s current crawler article does not mention it. Third parties describe it, and Claude-Web, as deprecated; that status is not documented by Anthropic.
Rank SniperEvidence: verified
RankSniperBot
Fetches a handful of public storefront pages when someone requests a Rank Sniper scan of that store.
- User-Agent
RankSniperBot/1.0 (+https://ranksniperhq.com/bot)- robots.txt
- Honours robots.txt, per the operator.It reads robots.txt before anything else; a disallowed path is not fetched by any identity, and a refused catalogue ends the scan.
- IP ranges
- Not published.
- Rank Sniper
- This is the scanner itself; it is not scored.
The crawler behind the free scan. Two lines of robots.txt turn it off.
Where Shopify’s edge refuses a declared request, it retries that single request with a standard browser profile and records the fallback in the scan trace.