Skip to content

Crawler rule inspection

Free robots.txt checker

Test robots rules against a specific URL. Inspect selected user-agent groups, matching lines and the reason a path is allowed or blocked.

Use the Robots.txt checker

Your input

Robots.txt checker inputs

Enter the exact page URL to test, including its path and query string. We fetch this origin’s /robots.txt unless you paste rules below.

If provided, we analyze this text against the URL instead of fetching robots.txt.

0 / 16,000

No account required. Fair-use limits apply. This check analyses the input you provide.

Check the details. Then check the answers.

See when AI names your brand and which pages it cites in a free visibility report.

Get your free report

Test the rules against a specific path

  1. Enter the URL the rules should match

    Include its path and query string. Paste rules to test an edit, or leave the text blank to fetch the origin’s /robots.txt.

  2. Inspect each crawler’s outcome

    The tool parses agent groups and sitemap declarations, then shows allowed, blocked or unresolved outcomes with matching rule evidence.

  3. Review changes before publishing

    Compare the matching lines with your intended policy. Test other important paths separately; pasted rules are never uploaded to your site.

Robots.txt controls crawling. It does not protect confidential content or guarantee that a URL is removed from search.

Test the rule you mean

A small path difference can change the decision.

Robots.txt is a set of instructions about crawling particular paths. A broad disallow and a more specific allow can coexist. A dedicated crawler group can change which rules apply. This checker exposes the matching evidence instead of reducing the file to a single site-wide verdict.

Enter the complete URL you want to test. You can let the tool fetch the origin’s robots.txt or paste a proposed file. The URL remains required in both cases because the evaluator needs a concrete path and query string to compare with the rules.

A practical workflow

How to use the Robots.txt checker

  • Enter a complete public URL

    Include the path and relevant query string you want to test. The URL is required even when you provide all the robots text yourself.

  • Fetch the file or paste a proposal

    Leave the text area empty to fetch robots.txt. Paste up to 16,000 characters to evaluate proposed rules instead of retrieving that file.

  • Review evidence before changing policy

    Read the selected group and matching line for each token. Test both intended exclusions and allowed exceptions, then verify the deployed file separately.

Example

The path and the matching rule decide the policy result

The policy blocks /private/ and makes a more specific exception for /private/public/. Follow the rule that matches the submitted path.

/private/public/guide

The path falls inside both rule prefixes. Compare it with /private/report, which only matches the broader restriction.

Matching details

Specificity matters more than line order alone.

The evaluator supports wildcard asterisks and an end-of-path dollar sign. Path matching is case-sensitive. It selects the most specific matching user-agent groups, combines equally specific groups and uses the longest matching path rule. Allow takes precedence at equal path specificity.

The table covers Googlebot, Bingbot, GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot. They are crawler product tokens with different purposes. The checker does not send a request pretending to be any of them.

A rule specimen

Follow the group, the path and the winning line.

Trace a short policy from crawler group to matching path, then test the URLs that matter on your site.

Give the evaluator the exact path.

A permission result belongs to the URL tested. Path case and query parameters can affect matching, so avoid replacing a detailed address with the homepage.

Example target URLExample
  1. Originhttps://example.com
  2. Target path/private/public/guide
  3. Comparison path/private/report

These two paths can produce different decisions under the same file.

What the file cannot do

Crawl instructions are not access controls.

Do not put confidential information on a publicly reachable URL and rely on Disallow for protection. A robots rule is voluntary crawl policy. It does not authenticate visitors, secure data or guarantee removal from a search index.

Also distinguish a missing file from an unresolved retrieval. The tool treats 4xx responses other than 429 as an unavailable robots file permitting crawling. Network errors, 5xx responses and 429 remain Unknown, so a temporary failure does not become a fabricated clean policy.

AI crawlability checker

Check robots rules for named crawler tokens alongside page directives and source content.

Sitemap checker

Check sitemap XML, find duplicate locations and download the parsed inventory.

SEO analyzer & auditor

Inspect titles, headings, metadata, image alt attributes and indexing directives.

Questions, answered

Robots.txt checker: your questions answered

A closer look at the inputs, method and next steps.

sourcesignal.ai● ONLINE
Pick a question and I’ll explain how this tool works.

Why do I need a URL when I paste the file?

The rules are tested against a concrete path, including its query string. Pasting supplies the policy; the required URL supplies the target being evaluated.

Does pasting rules change my live robots.txt?

No. Pasted text is analyzed instead of fetching the robots file. Nothing is uploaded or published to your website, and the result describes only that proposed input.

Does the first matching line always win?

No. The evaluator selects applicable agent groups and the longest matching path rule. Allow wins a tie at equal specificity; simple first-line order does not describe this method.

Does Disallow remove a URL from Google?

Not necessarily. Crawl permission and indexing are different questions. Use the appropriate indexing controls and verification for your goal instead of treating robots.txt as a removal guarantee.

Are sitemap declarations validated here?

The parser reports valid sitemap directives it recognizes. It does not fetch their XML files. Use the sitemap checker to inspect a declared sitemap’s structure and entries.

Can it prove a crawler obeys the file?

No. It simulates policy matching for named tokens. Actual visits, cached policies, firewall responses and individual crawler behavior need separate server or provider evidence.

Does this inspect page-level noindex directives?

This robots-only tool does not fetch the page HTML for that purpose. Use the AI crawlability checker or SEO analyzer when you also need page-directive findings.

Navigate the AI landscape. Place your brand on the map.

See when AI names your brand, which competitors appear, and the pages those answers cite. Turn what you find into a practical next step.

Five AI engines. No card required.