Live public-access probe

Can an AI agent enter?

Fetch the public directives that named AI and search agents encounter before they can read a page. No login, script, or cached score.

Public pages onlyLive at request timeCheck citation signals
Directive traceidle
Ready.

Submit a public domain to replace this waiting state with fetch evidence and one row for every named agent the checker evaluates.

  • 01homepage responsewaiting
  • 02robots.txt ruleswaiting
  • 03meta + header directiveswaiting
no credentials · public directivesinstrument ready

Access is not authority.

What this result means

A readable page can still be unconvincing.

This tool proves only the public access declarations it can observe. It does not prove indexing, citation, recommendation, ranking, or a model’s future behavior.

Instrument protocol

One request. Four evidence steps.

The output stays close to what the server actually fetched and evaluated, so a policy decision never masquerades as a growth forecast.

  1. 01
    Fetch the public surface

    The checker requests the submitted homepage and robots.txt without asking for CMS, analytics, or hosting access.

    live request
  2. 02
    Resolve declared access

    It evaluates robots.txt, homepage meta-robots, and X-Robots-Tag directives against each named agent.

    deterministic
  3. 03
    Keep purposes separate

    Search, user-action, training, and advertising agents stay labeled so a policy choice is not mistaken for a visibility defect.

    purpose aware
  4. 04
    Hand evidence to the scan

    The full Revvye scan combines this access evidence with public-page structure, buyer-path, trust, mobile, and conversion checks.

    one input
How this check works

How to tell whether your website is blocking AI crawlers

A public access directive can stop an agent before page quality matters. This check reads robots.txt, homepage meta-robots, and X-Robots-Tag declarations for the 14 named agents in Revvye's detector, then reports exactly what it observed.

05 field notes06 decisions
01

What the checker actually reads

When you submit a domain, the tool fetches two things live: your homepage and your robots.txt. It does not use a cached index, and it does not need access to your hosting, your CMS, or your analytics. Everything it evaluates is public.

From those responses it evaluates three public directive surfaces for each named agent. These surfaces behave differently:

  • robots.txt — a Disallow rule naming the bot, or a blanket rule that catches it. This is the most common cause and the easiest to fix.
  • meta robots — a noindex tag in the HTML of the page itself. The crawler is allowed to fetch the page but told not to use it.
  • X-Robots-Tag — the same instruction delivered as an HTTP header rather than a tag. Invisible when you view the page source, which is why it survives so long.

A page can be perfectly written, fast, and well-structured and still be invisible, because access is decided before any of that is read.

02

The distinction that matters most: training bots vs search bots

A site owner may reasonably choose not to allow training collection while still allowing search or user-initiated retrieval. Those are different policy choices, so the checker keeps the detector's purpose label beside every agent.

The purpose label describes how the Revvye detector classifies the agent: training, search, user-initiated, or advertising. A detected block closes that declared access path; it does not by itself prove what any external product indexed, cited, or displayed.

The table below is generated from the same named-agent catalog used by the detector. For current policy consequences beyond the observed directive, confirm the agent's purpose in the operator's own documentation before changing access rules.

Named agentOperatorWhat Revvye reports
OAI-SearchBotOpenAIClassified by the Revvye detector as search; the result reports only whether a public directive blocks that named agent.
GPTBotOpenAIClassified by the Revvye detector as training; the result reports only whether a public directive blocks that named agent.
ChatGPT-UserOpenAIClassified by the Revvye detector as user-initiated fetch; the result reports only whether a public directive blocks that named agent.
ClaudeBotAnthropicClassified by the Revvye detector as search; the result reports only whether a public directive blocks that named agent.
anthropic-aiAnthropicClassified by the Revvye detector as training; the result reports only whether a public directive blocks that named agent.
PerplexityBotPerplexityClassified by the Revvye detector as search; the result reports only whether a public directive blocks that named agent.
Perplexity-UserPerplexityClassified by the Revvye detector as user-initiated fetch; the result reports only whether a public directive blocks that named agent.
Google-ExtendedGoogleClassified by the Revvye detector as training; the result reports only whether a public directive blocks that named agent.
GooglebotGoogleClassified by the Revvye detector as search; the result reports only whether a public directive blocks that named agent.
BingbotMicrosoftClassified by the Revvye detector as search; the result reports only whether a public directive blocks that named agent.
AdsBot-GoogleGoogleClassified by the Revvye detector as advertising; the result reports only whether a public directive blocks that named agent.
DuckAssistBotDuckDuckGoClassified by the Revvye detector as search; the result reports only whether a public directive blocks that named agent.
MistralAI-UserMistralClassified by the Revvye detector as user-initiated fetch; the result reports only whether a public directive blocks that named agent.
YouBotYou.comClassified by the Revvye detector as search; the result reports only whether a public directive blocks that named agent.
03

Reading your result

The headline number is how many of the 14 named agents were blocked by a detected public directive. Each row preserves the agent, operator, purpose classification, observed source, and allowed/blocked state returned by the live API.

  • Zero blocked — no tested public directive denied the named agent set in this fetch. That does not prove indexing, citation, or recommendation.
  • Only training agents blocked — this may reflect an intentional content-use policy. Confirm the decision with the site owner and the current operator documentation.
  • Search or user-initiated agents blocked — review the matched public directive and decide whether that access path should be closed.
  • No robots.txt at all — crawlers default to allowed, so this is not an error. It does mean you have no way to express a preference later without adding the file.

If the homepage fetch itself fails, the meta and header checks are limited — the tool says so explicitly rather than reporting a clean pass it cannot support. robots.txt rules are still evaluated in that case, because robots.txt is fetched separately.

04

Fixing a block you did not intend

robots.txt lives at the root of your domain and is plain text. Whoever maintains your site can edit it in a few minutes. The change is to remove or narrow the Disallow rule under the user-agent you want to admit. Longer, more specific rules win over broader ones, so you can open a single path without opening the whole site.

A meta robots noindex is edited in the page template or a page-level publishing setting. An X-Robots-Tag is delivered at the server or CDN layer. The live tool does not impersonate every named agent at a network edge, so agent-specific WAF behavior remains outside this result and needs separate infrastructure inspection.

Change one thing at a time and re-run the check. When the public fetch succeeds, the checker immediately repeats the same directive evaluation and exposes the new result.

05

What this tool does not tell you

Being readable is a precondition for being cited, not a cause of it. This check confirms the door is open. It does not tell you whether an assistant currently mentions your business, whether your pages state what you do in a form worth quoting, or how you compare to the competitor that is being cited instead.

It also covers one part of the 8-category public scan. The published taxonomy currently contains 43 checks across crawler access, structured data, markup, metadata, privacy, security, email posture, and technology signals. The scanner starts from public pages and moves private results into the email-linked flow.

Q/A

Common questions

01How do I know if my site is blocking GPTBot?

Enter your domain above. The result lists GPTBot by name, preserves its training-purpose classification from the detector catalog, and reports the public directive that allowed or blocked it. Confirm broader product effects in OpenAI's current documentation before changing policy.

02Does blocking AI crawlers hurt my Google rankings?

The checker keeps Google-Extended and Googlebot as separate named agents with separate purpose labels. It reports the public directive state; it does not measure search ranking impact. Confirm current product consequences in Google's documentation before changing either rule.

03I have no robots.txt. Is that a problem?

No. Crawlers treat a missing robots.txt as permission to read everything, so an absent file is a default-allow state. The only downside is that you have no mechanism to express a preference until you add one.

04Why does the checker say a bot is blocked when my robots.txt looks fine?

Because robots.txt is one of three directive surfaces this tool reads. A homepage noindex meta tag or X-Robots-Tag header can produce a blocked result even when robots.txt allows the agent. Agent-specific firewall behavior is outside this public result.

05Does this check whether ChatGPT actually mentions my business?

No. This tool checks access — whether the crawlers are permitted to read your pages. Whether an assistant chooses to mention or cite you is a separate question that depends on how legible and specific your content is once it has been read.

06Is the checker free, and does it need my email?

It is free and it does not ask for an email, an account, or a card. It fetches two public documents from your domain and reports what they say.

Continue the investigation

Keep the evidence moving.

01AI Citation Readiness CheckerAccess is step one. This checks the structured-data and metadata signals that can support clear machine interpretation.02Run the full free scanAll 43 published checks across 8 categories, not just crawler access.03The full rule taxonomyEvery rule the scanner evaluates, published — including how each one is detected.