Search access and model-training permissions are different controls. Choose each deliberately.
- OAI-SearchBot handles OpenAI search discovery; GPTBot has a separate training role.
- Google Search AI features use Googlebot controls, while Google-Extended governs separate uses.
- Robots rules do not prove firewall access, indexing or inclusion in a generated answer.
- Audit your public rules and verify provider documentation before changing production settings.
Keep asking about it in
You can often allow AI search discovery while restricting model-training use. The crawler tokens and controls differ by provider. Blocking every token with “AI” in its description can restrict the search access you actually want.
The AI visibility checker reports search and training controls separately. Only the search-root check contributes to its technical readiness score. A decision to restrict training is a publishing choice and earns no penalty in this tool.
Know which control you are changing
| Provider | Search or retrieval control | Separate training or content-use control |
|---|---|---|
| OpenAI | OAI-SearchBot; user-initiated requests also have a distinct ChatGPT-User agent | GPTBot |
| Anthropic | Claude-SearchBot; Claude-User handles user-directed retrieval | ClaudeBot |
| Perplexity | PerplexityBot; Perplexity-User handles user-directed requests | Read its current policy rather than assuming every bot is a training crawler |
| Googlebot and Google Search indexing/snippet controls for Search AI features | Google-Extended controls separate Gemini-related uses described by Google |
These are distinct purposes, not promises that every request observes identical rules. In particular, user-directed retrieval can differ from automated crawling. Check the provider’s current explanation before writing a blanket rule: OpenAI’s crawler documentation, Anthropic’s crawler policy, Perplexity’s crawler documentation and Google’s common crawlers.
Example: permit OpenAI search while restricting GPTBot
For a publisher who has chosen that specific policy, the relevant groups could be:
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
This is an illustration, not a complete robots.txt file to paste over your current one. Preserve your sitemap and other rules. Inspect existing groups for these agents, path-specific rules and wildcard rules before editing. Robots matching depends on the applicable user-agent group and matching path; an isolated line can be misleading.
The example does not give OpenAI permission to log in, bypass a firewall or access private content. Robots.txt is a crawler policy, not an access-control system. Keep authentication and server controls on private pages.
Google Search AI features use Search foundations
Google’s AI features documentation says pages need the normal Search foundations, including being indexed and eligible for a snippet. There is no special AI markup or AI text file requirement. Do not remove a legitimate noindex directive from a private or low-value page just to pass a checker.
Google-Extended is a separate control from Googlebot. Read its precise current scope in the official crawler documentation before choosing a policy. Allowing it is not a requirement for a better readiness score in our tool.
Verify more than the text file
Open your public robots.txt URL and confirm it returns the intended text, rather than a challenge page or server error. Then check the actual page you want discovered:
- Does it return useful public content with a successful HTTP response?
- Do relevant robots groups permit that specific path?
- Do its meta robots or HTTP headers block indexing or snippets?
- Is a firewall challenging requests from the provider’s documented infrastructure?
- For Google, what does Search Console’s URL Inspection show?
Our automated scan checks blanket root restrictions and the returned public homepage. It does not impersonate each crawler’s verified network, render every JavaScript application or audit every path. A passing robots check therefore means “no blanket root block found,” not “all AI platforms can access and recommend this business.”
Separate access from visibility
Access is only one prerequisite. The content still needs to explain the business accurately, and an answer engine must decide it is relevant to the question. A readable page can be absent from an answer; a business can be mentioned through third-party sources even when its own website is not cited.
Use the readiness checker to inspect foundations, then follow the manual AI recommendation test to record actual answers. Keep both kinds of evidence so you know whether you are fixing a technical restriction, correcting business information or improving useful content.
Policy references reviewed 24 September 2026. Provider controls can change; recheck the linked documentation when changing your site’s policy.