What it does
The robots.txt generator writes the robots.txt file that tells crawlers which parts of your site they may fetch. Choose the user-agents for the main rule group, list the paths to disallow and the exceptions to allow, add an optional crawl delay and one or more sitemap URLs, and the file updates as you type.
Presets cover the common cases in one click: Allow all, Block all (useful for staging sites) and a WordPress starting point. A separate setting blocks AI crawlers either for model training only or entirely, including AI search and assistant fetchers, and you can block any other bot such as SEO crawlers by name. Paths, user-agent tokens and sitemap URLs are checked, so a missing leading slash or a relative sitemap is reported instead of silently producing a broken file.
How to use
- Start from a preset or keep the default, which allows everything.
- In Disallow paths, list folders or patterns to keep crawlers out of, separated by commas, such as
/admin/, /cart/, /*?sort=. Use Allow paths for exceptions inside them. - Add your sitemap URL so search engines find it without being told separately.
- Pick how to treat AI crawlers and add any other bots to block.
- Download the file as
robots.txtand upload it to the root of your domain, so it is served athttps://example.com/robots.txt. Each subdomain needs its own file.
Example
Disallowing /admin/ with an exception for /admin/public/, adding a sitemap and blocking AI training crawlers produces:
# All crawlers
User-agent: *
Disallow: /admin/
Allow: /admin/public/
# AI training crawlers
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /
Sitemap: https://example.com/sitemap.xml
(The AI group is shortened here; the tool lists every crawler.) Several User-agent lines followed by rules form one group that applies to all of them, and a crawler follows only the most specific group that names it.
Tips
*matches any sequence of characters and$anchors the end of the URL, so/*.pdf$blocks every PDF.- When Allow and Disallow both match, Google applies the longest (most specific) rule.
- Don’t block CSS and JavaScript files that pages need to render, or search engines may misjudge your pages.
- Test changes with the robots.txt report in Google Search Console before relying on them.
FAQ
› Does Disallow in robots.txt remove a page from Google?
No. Disallow stops well-behaved crawlers from fetching a URL, but the URL can still be indexed (without a snippet) if other sites link to it. To keep a page out of search results, let it be crawled and add a noindex robots meta tag or X-Robots-Tag header instead.
› Which AI crawlers does the tool block?
Block training crawlers adds GPTBot, ClaudeBot, anthropic-ai, Google-Extended, Applebot-Extended, CCBot, Bytespider, Meta-ExternalAgent, Amazonbot and other bots that collect data for model training. Block all AI bots also adds AI search indexers and assistant fetchers such as OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User, which can stop your pages from being cited in AI answers.
› Do all crawlers obey robots.txt?
Reputable search engines and the major AI companies say they honour it, but robots.txt is a convention, not access control. Scrapers can ignore it, so protect private content with authentication and use server rules or a firewall for abusive bots.
› Is Crawl-delay supported?
Bing, Yandex and some other crawlers respect Crawl-delay; Google ignores it and adjusts its crawl rate automatically based on how fast your server responds.