What robots.txt does and does not do
robots.txt is a request to well-behaved crawlers about which paths to fetch. Search engines honour it. It is not access control - it does not prevent anyone from visiting a URL, and hostile scrapers ignore it entirely. Anything that must not be public needs authentication, not a Disallow line.
There is an unhelpful consequence worth understanding: a page blocked in robots.txt can still appear in search results, listed without a description, because the crawler was told not to fetch it but still saw links pointing to it. To keep a page out of the index you have to let the crawler fetch it and find a noindex directive - blocking it in robots.txt actually prevents that directive from ever being read.
The rules that cost sites their traffic
A single stray Disallow line can remove a whole site from search. The classic disaster is a staging configuration promoted to production with Disallow set to slash, blocking everything - and because nothing errors, it is often discovered weeks later when traffic has already collapsed.
The subtler version blocks CSS or JavaScript directories. Search engines render pages to evaluate them, and a page whose stylesheets are blocked renders as unstyled text, which can be judged as broken or as failing mobile usability. Blocking asset directories was once ordinary advice and is now actively harmful.
Declaring sitemaps
A Sitemap line in robots.txt tells any crawler where the sitemap lives without depending on it being submitted to a particular search console. It is a single line, it works for every search engine at once, and it is one of the cheapest things you can do for discoverability.
The URL in that line should be absolute, and it should point at a sitemap that returns 200 with URLs matching the canonical hostname. A sitemap listing URLs on a hostname that redirects is a common and quietly damaging mistake.
Frequently asked questions
- Does robots.txt keep a page out of Google?
- No. It stops the page being crawled, but a blocked URL can still be listed without a description. To exclude a page from the index, allow crawling and serve a noindex directive.
- Is robots.txt a security measure?
- No, and it works against you as one. It is a public file that lists the paths you want ignored, which is a map for anyone curious. Private content needs authentication.
- Should I block CSS and JavaScript?
- No. Search engines render pages to evaluate them, and blocked stylesheets or scripts make a page render as broken, which can damage how it is assessed.
- Where should I declare my sitemap?
- Add an absolute Sitemap line to robots.txt. Every crawler reads it, so it works without submitting to each search console individually.