Robots.txt Generator
Build a robots.txt file: crawl rules per user agent, the WordPress paths worth blocking, AI crawler rules and the sitemap line.
# robots.txt
# Crawl rules only. A blocked page can still be indexed from a link elsewhere,
# so use a noindex meta tag for pages that must stay out of results.
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /wp-login.php
Disallow: /xmlrpc.php
Disallow: /?s=
Disallow: /search/
Sitemap: https://example.com/wp-sitemap.xml
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
Build a robots.txt with the rules that matter: the WordPress paths worth blocking, the search pages that waste crawl budget, an AI crawler stance and the sitemap line.
How to use
- Start from “allow everything”. A robots.txt exists to make exceptions, not to gate the site.
- Keep the WordPress defaults. They block
/wp-admin/while leavingadmin-ajax.phpopen, which is exactly what WordPress serves by default. - Block internal search results. They are thin, infinite and the most common source of wasted crawling on a busy site.
- Decide the AI stance deliberately. Blocking training crawlers is different from blocking the ones that fetch a page to answer a question about it.
- Save the file at the domain root, as
/robots.txt. It cannot live in a subfolder and it does not apply across subdomains.
Example
The WordPress baseline with a sitemap:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /wp-login.php
Disallow: /?s=
Sitemap: https://example.com/wp-sitemap.xml
The Allow line wins over the Disallow above it because it is more specific. That is the rule the whole file depends on.
Pitfalls
- robots.txt controls crawling, not indexing. A blocked page linked from elsewhere can still appear in results, with no description.
- Blocking a page also stops the crawler seeing its
noindextag, so the two together achieve the opposite of what people expect. Disallow: /on a live site removes it from search within days, and recovery takes far longer. It belongs on staging only.- Paths are prefix matches.
Disallow: /cartalso blocks/cart-abandonedand/cartography. - Blocking
/wp-content/hides your CSS and images from the crawler, which makes the page render badly in the mobile friendliness test. Crawl-delayis ignored by Google entirely. Bing and Yandex honour it; Google’s rate is controlled in Search Console.- The file must be at the domain root and served as plain text. A subfolder copy is never read.
- Subdomains need their own file.
example.com/robots.txtsays nothing aboutshop.example.com. - AI crawler rules are voluntary. They record intent, and the crawlers that ignore robots.txt were never going to read them.
Compatibility
The original robots.txt specification dates to 1994 and was standardised as RFC 9309 in 2022. Allow, wildcards * and end anchors $ are supported by Google, Bing, Yandex and DuckDuckGo. Crawl-delay is not part of the standard: Bing and Yandex honour it, Google ignores it. AI crawler user agents change often, so revisit the list. The tool runs entirely in your browser.
Frequently asked questions
Will robots.txt keep a page out of Google?
noindex meta tag, and leave the page crawlable so the tag can be read.Should I block AI crawlers?
Where does the file go?
/robots.txt. WordPress serves a virtual one, so an uploaded file at the root overrides it.Does blocking /wp-admin/ hurt anything?
admin-ajax.php stays allowed. Plenty of front end features call it.