What is robots.txt?
Robots.txt is a plain text file at the root of a website, such as https://www.example.com/robots.txt. Google says it "tells search engine crawlers which URLs the crawler can access on your site."
A simple file looks like this:
User-agent: *
Disallow: /admin/
Sitemap: https://www.example.com/sitemap.xml
Each group names a crawler (User-agent) and the paths it may not visit. The Sitemap line points crawlers to the site's XML sitemap.
What can't robots.txt do?
- Keep a page out of Google. Google says robots.txt "is not a mechanism for keeping a web page out of Google." Use noindex or password protection instead.
- Force crawlers to comply. Google says "it's up to the crawler to obey them," and while reputable crawlers do, "other crawlers might not."
Why does robots.txt matter for a moving company?
One wrong line can block search engines from the whole site, for example a rule left over from a staging site. Robots.txt is also where a business allows or blocks specific AI crawlers, such as the ones ChatGPT and Perplexity use for search. See what are AI crawlers.