Your robots.txt file is a small text file at the root of your domain, for example https://example.com/robots.txt. It tells search engine crawlers which parts of your site they may and may not request. It is powerful, and because it is so simple, it is also easy to break. One wrong line can stop Google from crawling an entire website.
This article is part of our complete guide: Why Is My Website Not on Google? The Complete Indexing Guide.
In Google Search Console, problems show up in two ways: “Blocked by robots.txt” (Google did not crawl the page because robots.txt disallows it) and “Indexed, though blocked by robots.txt” (Google indexed the URL from links, without being able to read it). This guide covers both.
How robots.txt works in 60 seconds
A robots.txt file is made of groups. Each group starts with a User-agent line naming the crawler, followed by Disallow and Allow rules. For example:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://example.com/sitemap_index.xml
This tells every crawler to stay out of /wp-admin/, except the AJAX file that many themes need, and points them to the sitemap. Rules match the start of the URL path, so Disallow: /shop also blocks /shop-now/ and /shopping-guide/. That detail causes many accidental blocks.
Important things robots.txt does not do
- It does not remove pages from Google. A blocked URL can still appear in results if other pages link to it, usually with no description. To keep a page out of the index, use a noindex tag and allow crawling so Google can see it.
- It is not security. The file is public. Anyone can read it, and listing private folders there can actually advertise them. Protect private areas with passwords and proper permissions.
- It is not followed by every bot. Reputable search engines obey it; scrapers and malicious bots often ignore it.
Common causes of “Blocked by robots.txt”
1. A leftover “Disallow: /” from development
Developers often block the whole site during a build. If that file goes live, Google stops crawling everything. Look for a line that reads exactly Disallow: / under User-agent: *.
2. Rules that are broader than intended
A rule meant for one folder can catch other URLs because of prefix matching. Wildcards (*) and end-of-URL markers ($) make this even easier to get wrong.
3. Blocking CSS, JavaScript or image folders
Older advice suggested blocking /wp-includes/ or theme folders. Today Google renders pages like a browser, so blocking the files that style and power your pages can make Google see a broken layout and misjudge mobile-friendliness.
4. Blocking parameter URLs that are linked internally
Blocking ?add-to-cart= or ?orderby= is often sensible, but if important URLs carry parameters they will be blocked too.
5. Platform or plugin defaults
Some SEO and security plugins write robots.txt rules. Shopify, Wix and Squarespace generate their own files. A change to settings can alter what is blocked without you editing anything directly.
How to test your robots.txt
- Open
yourdomain.com/robots.txtin a browser and read it line by line. - In Search Console, open Settings › robots.txt to see the version Google last fetched, when it was fetched and whether there were errors.
- Use URL Inspection on a blocked URL. The result will state that crawling is blocked by robots.txt and, with the live test, confirm whether the current file still blocks it.
- Use a robots.txt testing tool to try specific URLs against your rules before you publish changes.
How to fix it on each platform
WordPress
WordPress serves a virtual robots.txt unless a physical file exists in the root folder. Rank Math and Yoast let you edit it from the dashboard. A sensible default for most small business sites is very short:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://yourdomain.com/sitemap_index.xml
Avoid blocking /wp-content/ or /wp-includes/. Also check Settings › Reading: “Discourage search engines” changes how WordPress treats indexing. Our WordPress services include a robots and sitemap review.
Wix
Wix generates robots.txt automatically and lets you edit it in the SEO settings under robots.txt editor. Only change it if you know why; resetting to default is a safe fix if something looks wrong. See our Wix SEO checklist.
Squarespace
Squarespace’s robots.txt cannot be edited. It blocks things like internal search and some system URLs by design. If an important page is blocked, it is usually because of its URL or because the page is in a disabled or hidden state.
Shopify
Shopify generates a default file, which can be customised with a robots.txt.liquid template. Default blocks for cart, checkout and filtered collection URLs are normally correct. Make sure custom edits do not block /collections/ or /products/.
Fixing “Indexed, though blocked by robots.txt”
This status means Google indexed a URL it was not allowed to crawl. Decide what you want:
- If the page should rank: remove the block so Google can crawl it properly.
- If the page should not appear in Google: temporarily allow crawling, add a noindex tag, wait for Google to recrawl and drop it, then re-block if needed. Blocking alone will not remove it.
After you fix it
Search Console normally refetches robots.txt within a day. You can ask for a recrawl from the robots.txt report. Then request indexing for important pages and click “Validate fix” in the Page indexing report. Pages blocked for a long time may take a few weeks to regain their previous visibility.
A robots.txt checklist
- No
Disallow: /underUser-agent: *on a live site. - CSS, JavaScript and images are crawlable.
- Only admin, cart, checkout, internal search and true duplicates are blocked.
- The sitemap URL is listed and correct.
- Private content is protected by passwords, not robots.txt.
- Pages you want out of Google use noindex, not Disallow.
Where this fits in a full audit
Robots.txt is the first gate in our technical SEO audit. If Google cannot crawl, nothing else matters. After that we check noindex tags, canonicals, sitemaps and content quality. Projects like M10 News and ClariCast show how crawl settings and content work together.
Related guides and services
- Excluded by noindex tag: how to find and remove an accidental noindex.
- Website not showing on Google: the full diagnosis for sites that are missing from search.
- Discovered – currently not indexed: why Google postpones crawling and how to speed it up.
- Soft 404 errors: why Google treats some live pages as missing.
- Google indexing fix service: we diagnose and fix pages Google will not index.
Get expert help
If you would rather spend your time running your business, MIVAQ can handle this for you. We audit the site, fix the root causes, document every change and show you the before-and-after in Search Console, so you know exactly what improved and why. Start with our Google indexing fix service, browse real client results in our case studies, or contact us for a free, no-pressure review of your website.
Frequently asked questions
Can robots.txt hurt my SEO?
Yes, if it blocks pages, CSS or JavaScript that Google needs. It cannot help a page rank, it can only stop crawling.
Should I block AI crawlers in robots.txt?
That is a business decision. You can add groups for specific AI user agents. It does not affect Google Search rankings.
Where is robots.txt on WordPress?
At yourdomain.com/robots.txt. It is usually virtual, generated by WordPress or your SEO plugin, unless a physical file exists in the root folder.
How do I add my sitemap to robots.txt?
Add a line such as “Sitemap: https://yourdomain.com/sitemap_index.xml”. Also submit it in Search Console.