Skip to content

Can AI crawlers reach your site.

Per OpenAI's crawler page (read October 6, 2026), blocking OAI-SearchBot keeps a site out of ChatGPT search answers and blocking GPTBot does not. Paste your robots.txt to see which of 12 crawlers it blocks.

Check your robots.txt

Use this checker if a host, plugin or site-builder setting may block AI crawlers. The checker reads a pasted robots.txt by RFC 9309 rules. It also checks that your phone number and prices are in a page's raw HTML. Vercel's study, published December 17, 2024, found that no OpenAI, Anthropic or Perplexity crawler it measured ran JavaScript.

Copy all of yoursite.com/robots.txt, or tick the box below if it shows not found. Nothing is uploaded.

On the page, choose View page source and copy it. Do not use Inspect.

Examples

How it decides

To check a verdict yourself, apply rules 1 to 8 to your file and compare with the worked examples.

The rules this checker applies, in order, with the source of each
StepRuleWhere it comes from
1. InputThe pasted text is read as plain text, up to 500 KiB. The path must start with /, or be a full address (its path is used). A box ticked "no robots.txt file" means no rules.RFC 9309 section 2.5 sets the minimum parsing limit at 500 KiB. Google documents the same limit.
2. LinesText after # is a comment. A line is Field: value. Only User-agent, Allow and Disallow are used. Other lines, such as Sitemap and Crawl-delay, are skipped and do not end a group. A rule before the first User-agent line is ignored.RFC 9309 sections 2.2.1 to 2.2.4.
3. GroupsA group is one or more User-agent lines followed by Allow and Disallow lines. A User-agent line after a rule starts a new group.RFC 9309 section 2.1.
4. Which groupA crawler uses every group that names its token, ignoring case, merged into one. If no group names it, it uses the * group, and several * groups are merged the same way. If there is none, no rule applies and the path is allowed.RFC 9309 section 2.2.1. Merging several * groups is our reading.
5. Which ruleOf the rules in that group that match the path, the longest pattern wins. If an Allow and a Disallow tie on length, Allow wins. No matching rule means allowed. The path /robots.txt is always allowed.RFC 9309 section 2.2.2. Google measures length by the rule path.
6. PatternsA pattern matches from the first character of the path and is case-sensitive. * matches any run of characters. A final $ means the path must end there. A rule with no path is ignored, and a trailing * is ignored. A path that starts with *, as in RFC 9309's own example Disallow: *.gif$, is read as a pattern.RFC 9309 sections 2.2.2 and 2.2.3. The last two points come from Google.
7. Result per crawlerRules 4 to 6 give blocked or allowed, but only for crawlers whose vendor says they follow robots.txt. For ChatGPT-User and Perplexity-User the vendor says rules may not apply, and for OAI-AdsBot it says nothing. Their result is "No promise from the vendor", with what the file says shown beside it.The vendor pages in the crawler table, read October 6, 2026.
8. VerdictSearch crawlers blocked: any of OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, Bingbot is blocked. Other crawlers blocked: only a training or user-request crawler is blocked. Otherwise nothing is blocked.Ours. The search, training, training and grounding, user request and ads labels are our reading of each vendor's description.
9. Page textFrom the pasted HTML, remove comments and script, style and template blocks. Keep noscript, which a fetch without scripts shows. Remove tags, decode entities ($ is $), join text across inline tags such as span and b, collapse spaces and ignore case.Ours.
10. Each factIn the page text: found there. In markup: found only elsewhere in the HTML, such as a script, JSON-LD, a tel: link or a style block. Not in the raw HTML: found nowhere. A phone number matches with any spaces, dashes, dots or brackets between its digits. Other facts match whole words and numbers, so DC does not match inside abcdcba.Ours.
11. LimitsUp to 6 facts of 2 to 80 characters, and HTML up to 2,000,000 characters. Page text length is shown for information and is not scored. Anything else gets a message, not a result.Ours.
Worked examples to check by hand, computed by the same function the tool runs
robots.txtPathResult
User-agent: GPTBotDisallow: //GPTBot is blocked. OAI-SearchBot is allowed: no group names it and there is no * group (rule 4).
User-agent: *Disallow: //9 crawlers are blocked, including Googlebot and Bingbot. 3 have no promise from their vendors: ChatGPT-User, OAI-AdsBot, Perplexity-User (rule 7).
User-agent: *Disallow: /User-agent: OAI-SearchBotAllow: //OAI-SearchBot is allowed, because a group that names it replaces the * group (rule 4). The other 8 crawlers that follow robots.txt are blocked.
User-agent: *Disallow: /Allow: /pricing/pricingNo crawler is blocked. Allow: /pricing is 8 characters and Disallow: / is 1, so the longer rule wins (rule 5). On the path / the same file blocks 9 crawlers.
User-agent: *Disallow: /privateAllow: /private/private/menuNo crawler is blocked. The two rules are the same length, so Allow wins (rule 5).
User-agent: GPTBotDisallow: /*.pdf$/menu.pdf, then /menu.pdf?v=2GPTBot is blocked on /menu.pdf and allowed on /menu.pdf?v=2, because $ needs the path to end in .pdf (rule 6).

Crawlers it checks

These 12 crawlers come from five vendors' own pages, read October 6, 2026. The labels are ours.

The 12 crawlers, what each vendor says about them, and whether the vendor says they follow robots.txt
CrawlerWhat the vendor saysSource
OAI-SearchBotOpenAI, searchSurfaces websites in ChatGPT's search features. A site opted out of it is not shown in ChatGPT search answers, though it can still appear as navigational links. OpenAI says a change takes about 24 hours.Follows robots.txt, per the vendor: yes.OpenAI crawlers
GPTBotOpenAI, trainingCrawls content that may be used to train OpenAI's generative AI foundation models. Indicates the content should not be used to train those models. OpenAI says it is independent of OAI-SearchBot.Follows robots.txt, per the vendor: yes.OpenAI crawlers
ChatGPT-UserOpenAI, user requestVisits a page when a user asks ChatGPT or a custom GPT something. OpenAI says robots.txt rules may not apply, because the user starts the request, and that it is not used to decide whether content appears in Search.Follows robots.txt, per the vendor: no promise, user requests may ignore it.OpenAI crawlers
OAI-AdsBotOpenAI, ads checkChecks the safety of pages submitted as ads on ChatGPT, and visits only those pages. OpenAI's page does not say whether it follows robots.txt or what blocking it does.Follows robots.txt, per the vendor: not stated.OpenAI crawlers
PerplexityBotPerplexity, searchSurfaces and links websites in Perplexity search results. Perplexity says it is not used to crawl content for AI foundation models. Perplexity recommends allowing it so a site appears in its search results, and says a change may take up to 24 hours.Follows robots.txt, per the vendor: yes.Perplexity crawlers
Perplexity-UserPerplexity, user requestVisits a page when a user asks Perplexity a question, and links to it in the answer. Perplexity says that, because a user requested the fetch, it generally ignores robots.txt rules.Follows robots.txt, per the vendor: no promise, user requests may ignore it.Perplexity crawlers
Claude-SearchBotAnthropic, searchNavigates the web to improve the quality of search results for users. Anthropic says disabling it may reduce a site's visibility and accuracy in user search results.Follows robots.txt, per the vendor: yes.Anthropic help page
ClaudeBotAnthropic, trainingCollects web content that could contribute to training Anthropic's models. Signals that the site's future material should be excluded from Anthropic's training datasets.Follows robots.txt, per the vendor: yes.Anthropic help page
Claude-UserAnthropic, user requestVisits a page when a user asks Claude a question. Anthropic says disabling it stops Claude retrieving the content for a user's query, which may reduce visibility in user-directed web search.Follows robots.txt, per the vendor: yes.Anthropic help page
GooglebotGoogle, searchCrawls for Google Search, including Discover and all Search features. Google says robots.txt rules for Googlebot control how a site is crawled for Search. To be a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to show with a snippet.Follows robots.txt, per the vendor: yes.Google crawlers listGoogle AI features
Google-ExtendedGoogle, training and groundingA robots.txt token with no user agent of its own. It governs use of crawled content for training future Gemini models and for grounding in Gemini Apps. Google says it does not affect a site's inclusion in Google Search.Follows robots.txt, per the vendor: yes.Google crawlers list
BingbotMicrosoft, searchBing's crawler. Microsoft Advertising says Copilot is powered by Bing's search index. Microsoft says Bing respects the preferences a site expresses in robots.txt.Follows robots.txt, per the vendor: yes.Bing blog, June 17, 2025Bing blog, February 10, 2026Microsoft Advertising, October 8, 2025

What it cannot tell you

What this checker does not cover, with the source for each point
Not coveredWhat to knowSource
Firewalls, CDNs and bot protectionA firewall can block a crawler that robots.txt allows, and this checker never sees it. OpenAI recommends allowing OAI-SearchBot in robots.txt and allowing requests from its published IP ranges. Perplexity says a firewall may need its bots allow-listed. Anthropic says blocking its IP addresses can stop it reading your robots.txt.OpenAI, Perplexity and Anthropic, read October 6, 2026.
Crawlers that ignore robots.txtrobots.txt is a request, not a lock: RFC 9309 says its rules are "not a form of access authorization". OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them.RFC 9309, OpenAI and Perplexity.
Whether an assistant names youBeing allowed is not the same as being named. Google says a page that meets every requirement is still not guaranteed to be crawled, indexed or served. Microsoft Advertising says no "secret sauce" guarantees selection in AI answers. How assistants choose is in how AI assistants choose local businesses, and whether they send visitors is in how to see if AI sends you visitors. Our GEO page says what we do about it.Google AI features and Microsoft Advertising.
Indexingrobots.txt allowing Googlebot does not make a page indexed. Google says a page must be indexed and eligible to show with a snippet to be a supporting link in AI Overviews and AI Mode, with no other technical requirement. Google suggests verifying your site in Search Console to find technical issues.Google AI features, last updated 2025-12-10.
Which host the file belongs toA robots.txt applies only to the host, protocol and port that serve it. The file on www.example.com does not cover shop.example.com. Paste the file from the same host as the page you test.Google, last updated 2026-08-31.
Your file's status codeThe checker reads pasted text, not the status your server returns. RFC 9309 says a 4xx status means a crawler may access anything, and a 5xx status means it must assume everything is blocked. Google treats 4xx other than 429 as no file, and for the first 12 hours after a 5xx it stops crawling the site.RFC 9309 section 2.3.1 and Google, last updated 2026-08-31.
Whether crawlers run scripts todayThe facts check shows what the server sends. Vercel published its study on December 17, 2024. It found that no OpenAI, Anthropic or Perplexity crawler it measured ran JavaScript, and that Gemini uses Googlebot's rendering. We have not measured any crawler in 2026. Text hidden by CSS counts as page text here, because raw HTML does not show styles.Vercel. Vercel sells hosting, and our site is hosted there.
llms.txt and other AI text filesThis checker does not read them. Google says no AI text files or special markup are needed to appear in AI Overviews and AI Mode. What holds up is in llms.txt and GEO claims, checked.Google AI features.
A live test from outsideThe checker never fetches your site, so nothing you paste leaves your browser. To see what a crawler is served after your firewall, check your server or CDN logs for its user agent.Ours.

Sources

Sources for every external figure and rule on this page, with date and who publishes each
SourceWhat it saysDateWho publishes it
OpenAI, Overview of OpenAI crawlersFour user agents. OAI-SearchBot surfaces sites in ChatGPT search, and an opted-out site is not shown in its answers, though it can appear as navigational links. GPTBot is for training. ChatGPT-User acts on user requests and robots.txt rules may not apply. OAI-AdsBot checks pages submitted as ads. Search changes take about 24 hours.No date shown. Read October 6, 2026.OpenAI, which sells ChatGPT and ChatGPT Ads.
Perplexity, Perplexity CrawlersPerplexityBot surfaces and links sites in search results and is not used for foundation models. Perplexity-User generally ignores robots.txt. Changes may take up to 24 hours. A firewall may need its bots allow-listed.The page shows no date. Its metadata says modified January 29, 2026. Read October 6, 2026.Perplexity, which sells an AI search product and an API.
Anthropic, Does Anthropic crawl data from the webThree robots: ClaudeBot (training), Claude-User and Claude-SearchBot. They honor robots.txt directives. Blocking by IP address may not work and can stop the file being read.April 7, 2026. Read October 6, 2026.Anthropic, which sells Claude.
Google, List of Google's common crawlersGooglebot rules affect Google Search, including Discover and all Search features. Google-Extended is a robots.txt token for Gemini training and grounding in Gemini Apps, and does not affect inclusion in Google Search. Common crawlers always obey robots.txt rules.Last updated 2026-07-14 UTC. Read October 6, 2026.Google, which sells search ads and Gemini products.
Google, AI features and your websiteA page must be indexed and snippet-eligible to be a supporting link in AI Overviews or AI Mode. No AI text files or special markup are needed. Crawling should be allowed in robots.txt and by any CDN or hosting infrastructure. Indexing and serving are not guaranteed.Last updated 2025-12-10 UTC. Read October 6, 2026.Google, which sells search ads.
Google, How Google interprets the robots.txt specification500 KiB limit. The file applies only to its own host, protocol and port. A rule with no path is ignored. A trailing * is ignored. The longest rule path wins and the least restrictive rule wins a conflict. Handling of 4xx and 5xx statuses. Google generally caches robots.txt for up to 24 hours.Last updated 2026-08-31 UTC. Read October 6, 2026.Google, which sells search ads.
IETF, RFC 9309 Robots Exclusion ProtocolGroups, matching by product token, longest match, Allow winning a tie, * and $, status-code handling, a 24-hour cache limit and a 500 KiB minimum parsing limit. Says the rules are not access authorization.September 2022. Read October 6, 2026.The IETF, a standards body, which sells nothing. Three of the four authors list Google LLC.
Microsoft Bing blog, Start Using Bing Webmaster ToolsRefers to Bingbot visiting a site and lists a robots.txt tester that confirms Bingbot can access key content.June 17, 2025. Read October 6, 2026.Microsoft, which runs Bing and Copilot and sells Microsoft Advertising.
Microsoft Bing blog, AI Performance in Bing Webmaster ToolsSays Bing respects all content owner preferences expressed through robots.txt.February 10, 2026. Read October 6, 2026.Microsoft, which runs Bing and Copilot and sells Microsoft Advertising.
Microsoft Advertising, Optimizing your content for inclusion in AI search answersSays Copilot is powered by Bing's search index, and that no secret sauce guarantees selection in AI answers.October 8, 2025. Read October 6, 2026.Microsoft, which sells Microsoft Advertising.
Vercel, The rise of the AI crawlerPublished with MERJ from Vercel's network and nextjs.org logs. None of the major AI crawlers it measured rendered JavaScript, including OpenAI's OAI-SearchBot, ChatGPT-User and GPTBot, Anthropic's ClaudeBot and Perplexity's PerplexityBot. Gemini uses Googlebot's infrastructure and renders.December 17, 2024. Read October 6, 2026.Vercel, which sells hosting. Our site is hosted on Vercel.
Squarespace, Request that AI models exclude your siteThe Block known artificial intelligence crawlers box adds 26 named bots to robots.txt, including ClaudeBot, Google-Extended and GPTBot. It is off by default. The list does not name OAI-SearchBot, ChatGPT-User, PerplexityBot or Claude-SearchBot.Last updated January 8, 2026. Read October 6, 2026.Squarespace, which sells website hosting and building.
Montrelia, own robots.txt and pricing pagerobots.txt fetched with curl: one group, User-Agent: * with Allow: /. The /pricing page HTML fetched with curl has $2,000 in its page text by this checker's rules.Fetched October 6, 2026.Montrelia, which sells websites, SEO and ads. Our own measurement, on one day.

Questions people ask

Does my website block ChatGPT?

Your site blocks ChatGPT search if robots.txt disallows OAI-SearchBot for a page, or if a firewall stops its requests. With no group naming OAI-SearchBot and no User-agent: * group, robots.txt blocks nothing. Disallowing GPTBot only opts content out of training. ChatGPT-User, which acts on a user's request, may ignore robots.txt (OpenAI, read October 6, 2026).

Should I block GPTBot?

Blocking GPTBot is a choice about training use. OpenAI says GPTBot and OAI-SearchBot are independent settings, so blocking GPTBot does not remove a site from ChatGPT search. To appear in ChatGPT search, allow OAI-SearchBot. The checker treats a blocked training crawler as a choice, not a fault. Our own robots.txt blocks neither (read October 6, 2026).

Does Squarespace's AI crawler setting block ChatGPT search?

Its list does not name OAI-SearchBot. Squarespace says the "Block known artificial intelligence crawlers" box adds 26 named bots to robots.txt, including GPTBot, ClaudeBot and Google-Extended, but not OAI-SearchBot, ChatGPT-User, PerplexityBot or Claude-SearchBot (help page updated January 8, 2026, read October 6, 2026). Paste your live robots.txt to see what it wrote.

How long does a robots.txt change take to work?

OpenAI says about 24 hours for search results, Perplexity says up to 24 hours, and Google says it generally caches robots.txt for up to 24 hours. RFC 9309 says crawlers should not use a cached copy for more than 24 hours unless the file is unreachable. We found no figure on the Anthropic or Microsoft pages we read.

Last updated October 6, 2026

Talk to us for 15 minutes.