Skip to content

robots.txt and sitemap.xml Basics — What They Tell Crawlers, and Which One Wins

When a search engine crawler visits a site, it usually doesn’t jump straight to your articles. It first checks robots.txt to learn where it may go, and it reads sitemap.xml to learn what pages exist. Both files are “notes for crawlers,” so they get confused a lot — but they say different things and matter at different moments. The last few posts (hreflang, JSON-LD, OGP) were about how a page should be interpreted. This one is about the step before that: how a crawler gets to the page at all.

Note: crawlers are also called bots or spiders. Googlebot and Bingbot are the best-known examples. Each identifies itself with a User-Agent HTTP header, and robots.txt can give different rules to different agents.

In one line: robots.txt is a “keep out” sign, sitemap.xml is a floor map

robots.txt sitemap.xml
Says “Please don’t crawl these paths” “These URLs exist on this site”
Nature Restriction (subtractive) Discovery hint (additive)
Location Exactly one, at the host root (/robots.txt) Anywhere (announced via robots.txt or Search Console)
Format Plain text XML
Enforcement Well-behaved crawlers obey; nothing forces them Listing a URL guarantees neither crawling nor indexing

Neither file physically blocks access or promises a search listing. Both are requests.

Writing robots.txt: four directives are enough

  • User-agent: — which crawler the following rules apply to (* = all)
  • Disallow: — path prefixes not to crawl
  • Allow: — exceptions inside a disallowed area
  • Sitemap: — the full URL of a sitemap (applies globally, not per agent)

Here is the actual robots.txt at en.wpmm.jp, where this English blog lives:

User-agent: *
Allow: /

Disallow: /config.php
Disallow: /includes/

Sitemap: https://en.wpmm.jp/sitemap.xml
Sitemap: https://en.wpmm.jp/blog/wp-sitemap.xml

Everything is allowed except a config file and an internal directory. The two Sitemap: lines list the landing-page sitemap and the blog sitemap side by side — for a reason explained below.

When both Allow: / and Disallow: /includes/ match a path, major search engines apply the longest (most specific) matching rule. For /includes/foo.php, /includes/ beats /, so it is disallowed. Order in the file doesn’t decide it; match length does.

One robots.txt per host, at the root

The most common robots.txt mistake is location. Crawlers only fetch /robots.txt at the root of each scheme + host. https://wpmm.jp/robots.txt and https://en.wpmm.jp/robots.txt are separate files, and a robots.txt inside a subdirectory is simply never read.

This blog is a live example. WordPress normally serves a virtual robots.txt even with no physical file. But when WordPress is installed in a subdirectory, as it is here (/blog/), that virtual file ends up at /blog/robots.txt — a URL crawlers never request. In fact it returns WordPress’s 404 page:

curl -s -o /dev/null -w "%{http_code}\n" https://wpmm.jp/blog/robots.txt
# 404
curl -s -o /dev/null -w "%{http_code}\n" https://wpmm.jp/robots.txt
# 200

So with a subdirectory install, editing robots.txt from inside WordPress (settings or plugins) reaches no crawler at all. Rules and sitemap locations have to go into the physical robots.txt at the host root — which is why the blog’s wp-sitemap.xml is declared in the root file above.

Note: robots.txt’s own status code matters. A 404 means “no restrictions,” so everything is crawlable. Repeated 5xx errors, however, can make major search engines pause crawling the whole site for a while, because they can’t tell what is forbidden.

robots.txt is not a lock

  • It is public; anyone can read the paths you list
  • Obeying it is voluntary; some bots ignore it
  • A disallowed URL can still appear in results as a bare URL if other sites link to it

Using Disallow to “hide” sensitive pages can backfire by advertising them. Protect private content with authentication or server-side access control. Treat robots.txt as a way to keep crawlers from wasting time on pointless areas, nothing more.

Inside sitemap.xml: URLs and last-modified dates

A sitemap lists URLs in XML: a <loc> per <url>, optionally with <lastmod>.

<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://wpmm.jp/blog/safe-wordpress-maintenance-with-ssh-wpcli/</loc>
    <lastmod>2026-05-13T15:30:06+09:00</lastmod>
  </url>
</urlset>

The spec also defines <changefreq> and <priority>, but Google has said it ignores both. What matters in practice is an accurate <loc> and a <lastmod> that reflects real changes. Bumping <lastmod> every day for no reason only teaches crawlers not to trust it.

Larger sites split sitemaps by type and tie them together with a sitemap index (<sitemapindex>). WordPress has shipped this in core since 5.5: /wp-sitemap.xml returns an index pointing to per-type files such as wp-sitemap-posts-post-1.xml, and WordPress keeps them updated as you publish.

“Which one wins?” — don’t make them compete

If a URL is disallowed in robots.txt but listed in the sitemap, the crawler follows robots.txt and doesn’t fetch it. A sitemap announces URLs; it cannot override crawl permission. Search Console flags this as “Submitted URL blocked by robots.txt.”

The real lesson is to keep the two files from sending contradictory signals:

  • Pages you want in search: allowed in robots.txt, listed in the sitemap
  • Pages you want out of search: removed from the sitemap (and see the next point before blocking them)

Never combine noindex with Disallow

<meta name="robots" content="noindex"> says “don’t list this page” — but a crawler can only read it after fetching the page. Disallow the page in robots.txt and the crawler never sees the noindex, so the bare URL can linger in results.

  • Disallow = “don’t look inside” (crawl control)
  • noindex = “look, then don’t list it” (index control)

To reliably remove a page from search, don’t block it in robots.txt; return noindex from the page.

How this blog lines them up

The wpmm-blog theme marks 404s, search results, and category/tag archives as noindex via the wp_robots filter — thin, near-duplicate listing pages otherwise pile up as “discovered, not indexed” and burn crawl attention:

add_filter('wp_robots', function ($robots) {
    if (is_404() || is_search() || is_category() || is_tag() || is_tax()) {
        $robots['noindex'] = true;
        unset($robots['index']);
    }
    return $robots;
});

It then removes those same taxonomies from the core sitemap, so nothing is “submitted in the sitemap yet noindexed”:

add_filter('wp_sitemaps_taxonomies', function ($taxonomies) {
    unset($taxonomies['category']);
    unset($taxonomies['post_tag']);
    return $taxonomies;
});

Not blocked in robots.txt (so noindex can be read), not in the sitemap (no contradiction), noindex on the page — three settings that agree with each other.

Common pitfalls

  • Editing a subdirectory robots.txt: only the host root counts
  • Launching with Disallow: / still in place from development (also check WordPress’s “discourage search engines” setting)
  • Disallowing pages you meant to noindex
  • Sitemap URLs that differ from canonical URLs (http vs https, www, trailing slash)
  • A plugin hijacking the sitemap URL: an old sitemap plugin left active can make core’s /wp-sitemap.xml return empty output. We hit this at launch and fixed it by deactivating the plugin and regenerating rewrite rules
  • “Hiding” secret paths in robots.txt, which publishes them instead

Summary

robots.txt restricts crawling and lives once at the host root. sitemap.xml helps discovery, and its value lies in accurate URLs and honest lastmod dates. Because their jobs differ, the goal isn’t deciding which wins but keeping them consistent — especially: to keep a page out of search, don’t Disallow it; serve noindex and leave it out of the sitemap. Good search visibility starts with giving crawlers correct directions, before any on-page tuning.