Two small text files at the top of a site tell search engines where they may go and what is worth reading. Neither is required, and a site with neither still gets indexed. Having both makes crawling more predictable, and their absence is one of the first things reviewers and audit tools note.
robots.txt: where crawlers may go
The file must be named robots.txt and sit at the root of the domain, so it loads at https://yourdomain.com/robots.txt. A crawler reads it before fetching anything else. A typical small-site file:
User-agent: *
Disallow: /admin/
Disallow: /staging/
Sitemap: https://yourdomain.com/sitemap.xml
User-agent: * means the rules apply to every crawler. Each Disallow line names a path they should not fetch. The Sitemap line tells them where your list of pages is. To allow everything, leave Disallow: empty.
Three things robots.txt does not do
- It is not security. The file is public, and it is a request, not a lock. Well-behaved crawlers honor it; anything else can ignore it, and a list of "disallowed" paths tells a curious person exactly where to look. Protect private folders with a password.
- It does not reliably remove a page from search results. A blocked page can still be listed, without a description, if other sites link to it. To keep a page out of results, let crawlers fetch it and put
<meta name="robots" content="noindex">in its head. The crawler has to be able to read the page to see that instruction. - It should not block your stylesheets and scripts. Search engines render pages as a browser would. If they cannot load the files that lay the page out, they may judge it broken on phones.
sitemap.xml: what is worth reading
A sitemap is a list of the pages you want indexed. On a small site with good internal links, crawlers would find them anyway; the sitemap gets new pages noticed sooner and gives you a report of which ones were accepted. The format:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://yourdomain.com/</loc>
<lastmod>2026-10-01</lastmod>
</url>
<url>
<loc>https://yourdomain.com/about.html</loc>
<lastmod>2026-09-14</lastmod>
</url>
</urlset>
Rules that matter in practice:
- Use full addresses, in the one form your site actually serves: the same
httpsand the samewwwor non-wwwchoice as the pages themselves. - List only pages you want in search results. Leave out sign-in screens, admin areas, search results and anything marked noindex.
lastmodis optional. If you include it, keep it truthful; a date that changes every day on a page that does not teaches crawlers to ignore it.- One file can hold up to 50,000 addresses, far more than a small site needs.
WordPress has generated a sitemap by itself since version 5.5, at /wp-sitemap.xml, and SEO plugins replace it with their own. On a hand-built site, write the file by hand and update it when you add a page.
Tell the search engines
The Sitemap line in robots.txt is enough for crawlers to find it. Submitting the address in Google Search Console and Bing Webmaster Tools as well gives you something more useful: a report showing which pages were indexed and the reason for any that were not.
Check your work
Open both addresses in a browser. Each should load as plain text or XML with a 200 status, not a 403 or 404 and not a redirect to the home page. If robots.txt returns 403, the cause is usually file permissions; set the file to 644.
Common questions
Do I need a robots.txt at all?
No. If the file is missing, crawlers assume everything is allowed. The reasons to have one are to point at the sitemap and to keep crawlers out of areas that waste their time, such as internal search results.
How long until a new page appears in search?
From hours to weeks. A sitemap and links from pages that are already indexed are the two things that speed it up. Search Console's URL inspection tool lets you ask for a specific page to be crawled.
Should images and PDFs go in the sitemap?
List the pages. Crawlers find the images on a page by reading the page. A PDF that is worth finding on its own can be listed like any other address.
Can one robots.txt cover my subdomains?
No. Each hostname has its own: blog.yourdomain.com/robots.txt is separate from the main site's file.