✦ Free SEO tool

Robots.txt Checker

Fetch and validate the robots.txt of any site, test your URL against its rules, and see at a glance which crawlers are blocked.

Fetching and analysing the page…

5 of 5 free checks left today. Create a free account for unlimited checks.

Robots.txt Checker - results and guide

What the Robots.txt Checker checks

  • Whether robots.txt exists, returns 200 and is served as plain text
  • Syntax: unknown directives, rules outside a group, paths without a slash
  • A site-wide block for Googlebot or Bingbot
  • Whether the URL you entered may be crawled, and which rule decides it
  • Blocked CSS and JavaScript paths that Google needs for rendering
  • Sitemap lines, crawl-delay and the 500 KiB size limit
  • Which AI crawlers (GPTBot, ClaudeBot, PerplexityBot and others) are blocked

What is robots.txt?

robots.txt is a plain text file at the root of a site, for example https://example.com/robots.txt. It tells crawlers which paths they may fetch. Rules are grouped by User-agent, and each group lists Disallow and Allow paths. A crawler follows the group that matches its name most specifically and falls back to the * group.

Within a group the longest matching rule wins, and Allow wins a tie. The characters * (any text) and $ (end of URL) work as wildcards. This tool applies the same rules Google documents.

What robots.txt can and cannot do

The file controls crawling, not indexing. A blocked URL can still appear in results, without a description, when other sites link to it. To keep a page out of search, allow crawling and use noindex, which you can verify with the Noindex Tag Checker.

The most damaging mistake is a leftover Disallow: / from a staging site, which blocks everything. The second is a robots.txt that returns a server error: Google then treats the whole site as off limits until the file works again. List your sitemap with a Sitemap: line and test it with the Sitemap Checker.

How to use the Robots.txt Checker

  1. Enter the domain or page address you want to check. You can leave out https://.
  2. Press Check robots.txt. The page is fetched from our servers, the way a search engine crawler would fetch it.
  3. Read the findings from top to bottom: problems first, then warnings. Each one says what to change.

Common problems and how to fix them

The whole site is blocked

Remove "Disallow: /" from the group that applies to search engines, or replace it with the specific folders you want to keep private.

robots.txt returns an HTML page

The file does not exist and the server answers with a normal page and status 200. Create a real robots.txt, or make missing files return 404.

CSS and JavaScript folders are disallowed

Old advice was to block them. Google now renders pages and needs those files. Remove the rules or add Allow lines for .css and .js.

Frequently asked questions

Where should robots.txt be located?
At the root of the host, such as https://example.com/robots.txt. Each subdomain and each protocol needs its own file.
Does robots.txt stop a page from being indexed?
No. It stops crawling. To prevent indexing, use a noindex robots meta tag or X-Robots-Tag header and leave the page crawlable.
Does Google obey crawl-delay?
No. Google ignores the Crawl-delay directive. Bing and Yandex respect it.
Should I block AI crawlers?
That is a business decision. Blocking GPTBot or ClaudeBot keeps your content out of AI training and answers, and also means those assistants cannot cite or link to you.
Monitor Today. Stay Ahead.

SEO gets visitors to your site. Uptime keeps them there.

Monitor 5 websites free forever and get alerted the moment one goes down. No credit card required.