Blocked by robots.txt: How to Read, Test and Fix Your Robots File

A single “Disallow” line can stop Google crawling your most important pages. Learn how robots.txt works, how to test it, and how to fix “Blocked by robots.txt” and “Indexed, though blocked by robots.txt”.
Blocked by robots.txt: How to Read, Test and Fix Your Robots File, MIVAQ guide cover

Your robots.txt file is a small text file at the root of your domain, for example https://example.com/robots.txt. It tells search engine crawlers which parts of your site they may and may not request. It is powerful, and because it is so simple, it is also easy to break. One wrong line can stop Google from crawling an entire website.

This article is part of our complete guide: Why Is My Website Not on Google? The Complete Indexing Guide.

In Google Search Console, problems show up in two ways: “Blocked by robots.txt” (Google did not crawl the page because robots.txt disallows it) and “Indexed, though blocked by robots.txt” (Google indexed the URL from links, without being able to read it). This guide covers both.

How robots.txt works in 60 seconds

A robots.txt file is made of groups. Each group starts with a User-agent line naming the crawler, followed by Disallow and Allow rules. For example:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/sitemap_index.xml

This tells every crawler to stay out of /wp-admin/, except the AJAX file that many themes need, and points them to the sitemap. Rules match the start of the URL path, so Disallow: /shop also blocks /shop-now/ and /shopping-guide/. That detail causes many accidental blocks.

Important things robots.txt does not do

  • It does not remove pages from Google. A blocked URL can still appear in results if other pages link to it, usually with no description. To keep a page out of the index, use a noindex tag and allow crawling so Google can see it.
  • It is not security. The file is public. Anyone can read it, and listing private folders there can actually advertise them. Protect private areas with passwords and proper permissions.
  • It is not followed by every bot. Reputable search engines obey it; scrapers and malicious bots often ignore it.

Common causes of “Blocked by robots.txt”

1. A leftover “Disallow: /” from development

Developers often block the whole site during a build. If that file goes live, Google stops crawling everything. Look for a line that reads exactly Disallow: / under User-agent: *.

2. Rules that are broader than intended

A rule meant for one folder can catch other URLs because of prefix matching. Wildcards (*) and end-of-URL markers ($) make this even easier to get wrong.

3. Blocking CSS, JavaScript or image folders

Older advice suggested blocking /wp-includes/ or theme folders. Today Google renders pages like a browser, so blocking the files that style and power your pages can make Google see a broken layout and misjudge mobile-friendliness.

4. Blocking parameter URLs that are linked internally

Blocking ?add-to-cart= or ?orderby= is often sensible, but if important URLs carry parameters they will be blocked too.

5. Platform or plugin defaults

Some SEO and security plugins write robots.txt rules. Shopify, Wix and Squarespace generate their own files. A change to settings can alter what is blocked without you editing anything directly.

How to test your robots.txt

  1. Open yourdomain.com/robots.txt in a browser and read it line by line.
  2. In Search Console, open Settings › robots.txt to see the version Google last fetched, when it was fetched and whether there were errors.
  3. Use URL Inspection on a blocked URL. The result will state that crawling is blocked by robots.txt and, with the live test, confirm whether the current file still blocks it.
  4. Use a robots.txt testing tool to try specific URLs against your rules before you publish changes.

How to fix it on each platform

WordPress

WordPress serves a virtual robots.txt unless a physical file exists in the root folder. Rank Math and Yoast let you edit it from the dashboard. A sensible default for most small business sites is very short:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://yourdomain.com/sitemap_index.xml

Avoid blocking /wp-content/ or /wp-includes/. Also check Settings › Reading: “Discourage search engines” changes how WordPress treats indexing. Our WordPress services include a robots and sitemap review.

Wix

Wix generates robots.txt automatically and lets you edit it in the SEO settings under robots.txt editor. Only change it if you know why; resetting to default is a safe fix if something looks wrong. See our Wix SEO checklist.

Squarespace

Squarespace’s robots.txt cannot be edited. It blocks things like internal search and some system URLs by design. If an important page is blocked, it is usually because of its URL or because the page is in a disabled or hidden state.

Shopify

Shopify generates a default file, which can be customised with a robots.txt.liquid template. Default blocks for cart, checkout and filtered collection URLs are normally correct. Make sure custom edits do not block /collections/ or /products/.

Fixing “Indexed, though blocked by robots.txt”

This status means Google indexed a URL it was not allowed to crawl. Decide what you want:

  • If the page should rank: remove the block so Google can crawl it properly.
  • If the page should not appear in Google: temporarily allow crawling, add a noindex tag, wait for Google to recrawl and drop it, then re-block if needed. Blocking alone will not remove it.

After you fix it

Search Console normally refetches robots.txt within a day. You can ask for a recrawl from the robots.txt report. Then request indexing for important pages and click “Validate fix” in the Page indexing report. Pages blocked for a long time may take a few weeks to regain their previous visibility.

A robots.txt checklist

  • No Disallow: / under User-agent: * on a live site.
  • CSS, JavaScript and images are crawlable.
  • Only admin, cart, checkout, internal search and true duplicates are blocked.
  • The sitemap URL is listed and correct.
  • Private content is protected by passwords, not robots.txt.
  • Pages you want out of Google use noindex, not Disallow.

Where this fits in a full audit

Robots.txt is the first gate in our technical SEO audit. If Google cannot crawl, nothing else matters. After that we check noindex tags, canonicals, sitemaps and content quality. Projects like M10 News and ClariCast show how crawl settings and content work together.

Related guides and services

Get expert help

If you would rather spend your time running your business, MIVAQ can handle this for you. We audit the site, fix the root causes, document every change and show you the before-and-after in Search Console, so you know exactly what improved and why. Start with our Google indexing fix service, browse real client results in our case studies, or contact us for a free, no-pressure review of your website.

Frequently asked questions

Can robots.txt hurt my SEO?

Yes, if it blocks pages, CSS or JavaScript that Google needs. It cannot help a page rank, it can only stop crawling.

Should I block AI crawlers in robots.txt?

That is a business decision. You can add groups for specific AI user agents. It does not affect Google Search rankings.

Where is robots.txt on WordPress?

At yourdomain.com/robots.txt. It is usually virtual, generated by WordPress or your SEO plugin, unless a physical file exists in the root folder.

How do I add my sitemap to robots.txt?

Add a line such as “Sitemap: https://yourdomain.com/sitemap_index.xml”. Also submit it in Search Console.

Keep reading

How to Start a Shopify Store – MIVAQ complete guide cover
ShopifyWeb
Step-by-step guide to starting a Shopify store in 2026: costs, plans, themes, products, payments, shipping, SEO and a full launch checklist.
Wix & Squarespace SEO – MIVAQ complete guide cover
SEOWeb
The complete 2026 guide to Wix SEO and Squarespace SEO: setup, titles, images, speed, local SEO, schema, redirects and fixing sites not
How to Build a WordPress Website – MIVAQ complete guide cover
WebWordPress
Step-by-step guide to building a WordPress website for your business in 2026: costs, hosting, themes, plugins, SEO, speed, security and a launch

Need help putting this into practice?

We turn ideas like these into working websites, stores and growth plans. Tell us what you are working on.