What Is robots.txt and How Does It Affect SEO?

Your website has a small text file that can help your SEO or quietly wreck it. It’s called robots.txt, and one wrong line in it can hide your whole site from Google.

So, what is robots.txt? It’s a plain text file that tells search engine crawlers and AI bots which parts of your website they can visit and which parts they should skip. It sits at the root of your domain, for example yourwebsite.com/robots.txt.

In this guide, you’ll learn how robots.txt works, how it affects SEO, what the rules mean, and how to avoid the mistakes that cost websites traffic. You’ll also see how it fits with AI crawlers like GPTBot, which matters more now that people find answers through ChatGPT, Perplexity, and Google AI Overviews.

What Is robots.txt?

robots.txt is a text file placed in the root folder of a website that gives instructions to web crawlers about which URLs they are allowed or not allowed to crawl. It follows a standard called the Robots Exclusion Protocol.

Think of it as a “house rules” sign at your front door. Polite visitors, like Googlebot and Bingbot, read the sign and follow it. Bad bots may ignore it.

Two things to remember:

  • robots.txt controls crawling, not indexing.
  • It’s a request, not a lock. It doesn’t protect private data.

How Does robots.txt Work?

Every time a crawler arrives at your site, it checks /robots.txt first, before visiting any other page. It reads the rules, then decides where to go.

The steps look like this:

  1. A crawler such as Googlebot lands on your domain.
  2. It requests yourwebsite.com/robots.txt.
  3. It looks for rules that match its name (its “user-agent”).
  4. It crawls the allowed pages and skips the blocked ones.

If your site has no robots.txt file, crawlers assume they can visit everything. That’s fine for many small sites, but larger sites usually need one.

What Does a robots.txt File Look Like?

Here is a simple robots.txt example:

User-agent: *

Disallow: /admin/

Disallow: /cart/

Allow: /admin/help/

Sitemap: https://www.yourwebsite.com/sitemap.xml

Here’s what each line does:

  • User-agent: names the crawler the rule applies to. * means all crawlers.
  • Disallow: tells the crawler not to visit that path.
  • Allow: lets the crawler visit a path, even inside a blocked folder.
  • Sitemap: shows crawlers where your XML sitemap is.

robots.txt Directives Explained

DirectiveWhat it doesExample
User-agentPicks which bot the rules apply toUser-agent: Googlebot
DisallowBlocks a URL or folder from crawlingDisallow: /private/
AllowPermits a URL inside a blocked folderAllow: /private/public-page/
SitemapPoints to your XML sitemapSitemap: https://site.com/sitemap.xml
Crawl-delayAsks bots to slow down (Google ignores it)Crawl-delay: 10

Two symbols are also useful:

  • * is a wildcard that matches any characters.
  • $ marks the end of a URL. For example, Disallow: /*.pdf$ blocks all PDF files.

How Does robots.txt Affect SEO?

robots.txt doesn’t directly boost your rankings. But it shapes how search engines spend their time on your site, and that affects your SEO in three main ways.

1. It helps manage your crawl budget

Search engines only spend a limited amount of time and resources crawling each site. This is called your crawl budget. If crawlers waste time on low-value pages like filters, internal search results, or cart pages, they may crawl your important pages less often. Blocking those low-value URLs helps crawlers focus on the pages that matter. If you want to understand this better, read our guide on what is crawl budget in SEO.

2. It reduces crawling of duplicate and junk pages

Many sites create hundreds of near-duplicate URLs through parameters, sorting options, and session IDs. robots.txt can stop crawlers from getting lost in them. For duplicate content you want consolidated, a canonical tag is often the better fix.

3. It protects your server

Aggressive bots can slow your site down. Blocking or limiting them keeps your site fast for real visitors, and speed is part of good SEO.

robots.txt vs noindex: What’s the Difference?

This is the most common point of confusion, so it’s worth getting right.

robots.txtnoindex tag
ControlsCrawlingIndexing
Where it livesA file at your site rootIn the page’s HTML or HTTP header
Stops page from appearing in Google?Not alwaysYes
Best forSaving crawl budget, blocking sectionsKeeping a page out of search results

Here’s the catch. If you block a page in robots.txt, Google can’t crawl it, so it can’t see a noindex tag on it either. The page may still show up in search results if other sites link to it, just without a description.

The simple rule: to keep a page out of Google, use noindex and leave the page crawlable. Use robots.txt to manage crawling, not to hide pages.

robots.txt and AI Crawlers

Search is no longer only Google. AI tools also crawl the web to learn from it and to find answers. Some common AI crawlers include:

  • GPTBot (OpenAI)
  • ChatGPT-User (OpenAI, for live browsing)
  • ClaudeBot (Anthropic)
  • PerplexityBot (Perplexity)
  • Google-Extended (controls whether Google uses your content for its AI models)
  • CCBot (Common Crawl)

You can allow or block each one in robots.txt. For example, to block GPTBot:

User-agent: GPTBot

Disallow: /

Should you block them? It depends on your goal. If you want your brand to be mentioned and cited in AI answers, blocking these crawlers can keep you out of the conversation. If you have paid or private content you don’t want used for training, blocking makes sense. Many businesses that want visibility choose to allow AI crawlers and focus on strong content and schema markup. Read what is AEO to see how AI answer engines pick sources.

How to Create a robots.txt File

Creating one takes only a few minutes.

  1. Open a plain text editor like Notepad. Don’t use a word processor.
  2. Write your rules using User-agent, Disallow, Allow, and Sitemap lines.
  3. Save the file as robots.txt (all lowercase).
  4. Upload it to your root directory so it loads at yourwebsite.com/robots.txt.
  5. Test it to make sure it works.

If you use WordPress, plugins like Yoast SEO, Rank Math, or All in One SEO let you edit robots.txt without touching any files. Shopify and Wix generate one automatically, and you can edit it in some plans.

Where Should robots.txt Be Located?

It must sit in the root of your domain. These will work:

  • https://www.example.com/robots.txt

These will not:

  • https://www.example.com/blog/robots.txt
  • https://www.example.com/Robots.txt (wrong case)

Each subdomain needs its own file. blog.example.com needs its own robots.txt, separate from example.com.

How to Test Your robots.txt File

Never publish changes without checking them. You have a few options:

  • Google Search Console: the robots.txt report shows the file Google found and any errors. Our guide on using Google Search Console covers the basics.
  • URL Inspection tool: it tells you if a specific URL is blocked by robots.txt.
  • Open the file in your browser: type your domain plus /robots.txt and read it.
  • Third-party robots.txt testers: many free SEO tools let you paste rules and test URLs.

Common robots.txt Mistakes That Hurt SEO

These errors show up again and again during SEO audits:

  1. Blocking the whole site. Disallow: / under User-agent: * stops all crawling. This often happens when a staging site’s robots.txt is copied to the live site.
  2. Blocking CSS and JavaScript files. Google needs these to render your pages properly. Blocking them can hurt how your pages are understood.
  3. Using robots.txt to hide private pages. The file is public. Anyone can read it, and it shows exactly which paths you’re trying to hide.
  4. Blocking a page and adding noindex. Google can’t see the noindex tag if it can’t crawl the page.
  5. Typos and wrong syntax. A misspelled directive is simply ignored.
  6. Forgetting the sitemap line. Adding your sitemap URL helps crawlers find your pages faster. Learn more in our guide on what a sitemap is and why your website needs one.
  7. Wrong file location or name. It must be robots.txt, in the root, in lowercase.

robots.txt Best Practices

  • Keep it short and simple.
  • Block only what you truly don’t want crawled, such as admin areas, cart, checkout, and internal search results.
  • Always include your XML sitemap URL.
  • Don’t block pages you want to rank.
  • Don’t block CSS, JS, or image files needed for rendering.
  • Test after every change.
  • Review the file regularly, especially after a redesign or migration.
  • Decide on purpose whether AI crawlers are allowed.

Real-World robots.txt Example for a WordPress Site

User-agent: *

Disallow: /wp-admin/

Allow: /wp-admin/admin-ajax.php

Disallow: /?s=

Disallow: /search/

Sitemap: https://www.yourwebsite.com/sitemap_index.xml

This blocks the admin area and internal search pages, keeps the AJAX file open (many themes need it), and points crawlers to the sitemap.

Frequently Asked Questions

What is robots.txt in simple words?

robots.txt is a small text file on your website that tells search engine bots and AI crawlers which pages they can visit and which they should avoid.

Is robots.txt necessary for SEO?

It’s not required, but it’s very helpful. Small sites can work without it. Larger sites, online stores, and sites with many filter or search URLs benefit from it because it protects crawl budget.

Does robots.txt stop a page from appearing in Google?

Not reliably. It stops crawling, but a blocked page can still be indexed if other sites link to it. Use a noindex tag to keep a page out of search results.

Can robots.txt improve rankings?

It doesn’t boost rankings directly. It helps by guiding crawlers to your important pages and keeping them away from low-value ones, which can lead to better crawling and indexing.

Where do I find my robots.txt file?

Type your domain followed by /robots.txt in your browser, for example yourwebsite.com/robots.txt.

Should I block AI crawlers like GPTBot in robots.txt?

It depends. Block them if you don’t want your content used for AI training. Allow them if you want your brand to appear in AI answers and citations.

What happens if I make a mistake in robots.txt?

A small mistake can block important pages or your whole site from being crawled, which can cause traffic to drop. Always test before publishing.

Is robots.txt case sensitive?

The file name must be lowercase (robots.txt). The paths inside it are case sensitive, so /Admin/ and /admin/ are treated as different.

Final Thoughts

robots.txt is small, but it has a big job. Used well, it helps search engines and AI tools spend their time on the pages that matter. Used badly, it can hide your best content.

Keep it simple, test every change, and remember the golden rule: robots.txt controls crawling, and noindex controls indexing. If you’d like expert help with technical SEO, crawl issues, or getting your brand cited in AI answers, check out the AI SEO services from BizClick Digital.

Leave a Reply

Your email address will not be published. Required fields are marked *