Short answer: paste your URLs and read the verdict on each. It names the robots rule, error or robots.txt block keeping a page out of Google, and flags a canonical that points elsewhere.
What noindex does, and how pages end up with it
A noindex rule tells Google not to show a page in results. It comes from a robots meta tag in the page’s <head> or from an X-Robots-Tag HTTP header, the only option for PDFs and images. Google’s robots meta tag documentation lists the rules it supports and says that when rules conflict, “the more restrictive rule applies”.
Pages usually get it by accident: a staging setting that shipped, a plugin option, a CDN rule adding a header everywhere. WordPress’s “Discourage search engines” box adds a noindex,nofollow robots meta tag to every page.
What this checker reads
- The response. The final status code and each redirect on the way. We judge the final URL: a 301 or 308 tells Google the destination is the canonical page, a 302 or 307 doesn’t. A page that returns a 4xx or 5xx has nothing to index.
- Robots meta tags.
robotsandgooglebottags, withnoneread asnoindex, nofollow. We flag tags outside<head>and rules Google doesn’t know, such asno-index. - X-Robots-Tag. Including rules for one crawler, like
googlebot: noindex. Other crawlers’ rules are shown but don’t block Google. - robots.txt. Whether it blocks Googlebot from the URL, with the rule and its line. Test other paths with the robots.txt tester.
- The canonical. A canonical to a different URL isn’t a block, so it shows as a warning. The canonical checker goes deeper.
The robots.txt trap
A Disallow in robots.txt stops Google crawling a page, so it never sees the noindex on it. Google says that “any information about indexing or serving rules will not be found and will therefore be ignored”, and that a blocked page can still appear in search results if other pages link to it. To get a page out of Google, let it be crawled and keep the noindex. We flag pages that have both.
What the verdicts mean
Can be indexed means nothing we can see stops it, not that Google has indexed it; the Google index checker looks that up. Can be indexed, but is a warning such as a canonical elsewhere. Blocked names the reason. After a fix, Google has to recrawl the page before anything changes; ask for that in the URL Inspection tool. Couldn’t check means the site refused our checker or didn’t answer.
Limits
We fetch as LaunchRankedBot, not Googlebot, so a server that treats Googlebot differently can show it something else. We read the HTML the server sends and don’t run JavaScript. Google says that when it sees noindex it may skip rendering and JavaScript execution, so a script that removes the tag may not work: put the right tag in the HTML you serve.