robots.txt and sitemap.xml Basics — What They Tell Crawlers, and Which One Wins
When a search engine crawler visits a site, it usually doesn’t jump straight to your articles. It first checks robots.txt to learn where it may go, and it reads sitemap.xml to learn what pages exist. Both files are “notes for crawlers,” so they get confused a lot — but they say different things and matter at different moments. The last few posts (hreflang, JSON-LD, OGP) were about how a page should be interpreted. This one is about the step before that: how a crawler gets to the page at all. Note: crawlers are also called bots or spiders. Googlebot and Bingbot are the best-known examples. Each identifies itself with a …