Skip to content
JustTools

Robots.txt Generator

Build a correct robots.txt in a minute — keep private paths out of search, choose which AI and SEO crawlers may visit, add your sitemap — then check any robots.txt and test URLs the way Googlebot reads them.

Looks up public records Free · no sign-up

At a glance

  • Generates robots.txt from simple choices: default access, paths to block with exceptions, an optional Crawl-delay and your sitemap URLs.
  • Lets you allow or block 26 named crawlers in four groups — search engines, AI search assistants (OAI-SearchBot, Claude-SearchBot, PerplexityBot), AI training bots (GPTBot, ClaudeBot, Google-Extended, CCBot …) and SEO crawlers (AhrefsBot, SemrushBot …).
  • Presets for WordPress, online shops, staging sites, and blocking AI training while staying in search and AI answers.
  • Validates any robots.txt with line numbers: rules outside a group, unsupported noindex, misspellings, relative sitemaps, blocked CSS and JavaScript, and Google’s 500 KiB limit.
  • Tests a URL against every crawler using Google’s matching rules (RFC 9309): the most specific user-agent group, the longest matching path, Allow winning ties — and shows the line that decided.
  • Loads a site’s live robots.txt for checking — directly from your browser when the site allows it, otherwise through Jina Reader. Building and testing happen on your device.

Step by step

How to create a robots.txt file

  1. 1

    Pick a starting point

    Choose a preset such as WordPress or Block AI training, or start from Allow everything.

  2. 2

    Add your rules

    List paths to keep out of search (like /admin/), allow or block specific crawlers, and add your sitemap URL.

  3. 3

    Test it

    Click Test a URL against this file to check important pages are still crawlable.

  4. 4

    Upload

    Download robots.txt and upload it to your site’s root so it opens at example.com/robots.txt.

Features

Everything you need, nothing you don’t

Correct by construction

Groups, paths and sitemap lines are written the way Google, Bing and AI crawlers expect, with full sitemap URLs.

AI crawler control

Opt out of AI training without disappearing from ChatGPT, Claude or Perplexity search answers — or block AI completely.

Smart presets

WordPress, online shop, staging and AI presets give a sensible starting point in one click.

Validator with line numbers

Finds the mistakes that silently break robots.txt — and jumps to the line.

URL tester

See which crawlers may fetch a page and which rule decides, before you upload the file.

Check any live site

Load a website’s robots.txt to audit it, or to start editing your own.

robots.txt rules at a glance

LineMeaningExample
User-agentWhich crawler the following rules are for; * means allUser-agent: GPTBot
DisallowPaths the crawler must not fetchDisallow: /admin/
AllowAn exception inside a disallowed pathAllow: /admin/help/
*Matches any charactersDisallow: /*?sort=
$Marks the end of the URLDisallow: /*.pdf$
SitemapFull URL of a sitemap, anywhere in the fileSitemap: https://example.com/sitemap.xml

What robots.txt can and can’t do

robots.txt controls crawling, not indexing. A blocked page can still appear in Google as a bare link if other sites link to it, and Google can’t see a noindex tag on a page it isn’t allowed to fetch. To keep a page out of search results, allow crawling and add <meta name="robots" content="noindex"> instead.

It is also public and voluntary: well-behaved crawlers such as Googlebot, Bingbot, GPTBot and ClaudeBot obey it, but it is not access control. Never list secret paths in it, and protect private areas with a login.

FAQ

Frequently asked questions

Where do I put robots.txt?

In the root of your domain, so it opens at https://example.com/robots.txt. Each subdomain (blog.example.com) needs its own file, and it must be plain text returning HTTP 200.

How do I block ChatGPT and other AI bots?

Add a group for each crawler with Disallow: /. To stop training only, block GPTBot, ClaudeBot, Google-Extended, CCBot and similar; to also leave AI search answers, block OAI-SearchBot, Claude-SearchBot and PerplexityBot. The “Block AI training” and “Block all AI” presets do this for you.

Will blocking Google-Extended remove my site from Google Search?

No. Google-Extended only controls whether Gemini may use your content for training and grounding. Search, including AI Overviews, uses Googlebot.

Does Google follow Crawl-delay?

No. Bing and Yandex do; Google adjusts its crawl rate automatically. If Googlebot overloads your server, return 503 or 429 temporarily.

Why does the validator say noindex isn’t supported?

Google stopped honouring Noindex lines in robots.txt in September 2019. Use a robots meta tag or the X-Robots-Tag HTTP header on the page instead.

Is blocking CSS and JavaScript a problem?

Yes. Google renders pages like a browser; without your CSS and JavaScript it may see a broken page and rank it lower. Keep folders such as /wp-includes/ and /_next/ crawlable.

Keep going

Browse every tool

Last updated Report a problem or suggest a feature