01

Choose tokens by the use they control

Read each provider's documentation separately. A token named in robots.txt is not necessarily a distinct HTTP User-Agent, and a search crawler is not automatically a training crawler. The table records the documented roles; an unspecified cell is a limit of the cited documentation, not proof that a provider has no other controls.

Provider and sourceSearch or answer surfacingTraining or other content-use controlUser-triggered fetching
GoogleGooglebot: crawl preferences affect Google Search, including its features. Common crawlers obey robots.txt during automatic crawling.Google-Extended: controls whether crawled content may be used for specified Gemini model training and grounding. It does not affect Google Search inclusion and is not used as a Google Search ranking signal.Google documents separate user-triggered fetchers; because a user requested the fetch, they ignore robots.txt rules.
BingBingbot: Bing's standard crawler. Its robots.txt documentation uses User-agent: Bingbot for a specific group.A separate training-control token is not stated in the cited documentation.A user-triggered token is not stated in the cited documentation.
OpenAIOAI-SearchBot: surfaces websites in ChatGPT search results. Sites that opt out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. Its robots.txt setting is independent of GPTBot.GPTBot: crawls content that may be used for foundation-model training; disallowing it indicates that site content should not be used for that training.ChatGPT-User: certain ChatGPT and Custom GPT user actions; robots.txt rules may not apply. It does not determine Search appearance.
PerplexityPerplexityBot: surfaces and links websites in Perplexity search results; it is not used to crawl content for AI foundation models.No separate training-control token is stated on this page.Perplexity-User: supports user requests and might fetch a page for an answer and link; it is not used for web crawling or to collect content for training AI foundation models. Perplexity lists Perplexity-User among the robots.txt tags webmasters can use. Perplexity-User controls which sites these user requests can access. It generally ignores robots.txt.

Google-Extended has no separate HTTP request User-Agent string: existing Google agents perform the crawling, while this token provides a control. Its grounding scope includes Gemini Apps and Grounding with Google Search on Vertex AI. Treating it as a training-only switch would conceal part of the documented effect.

02

Restrict training and grounding uses without naming search crawlers

For a site that intends to keep search crawling available, the following is an illustrative fragment, to be reviewed against its existing robots.txt:

```text User-agent: GPTBot Disallow: /

User-agent: Google-Extended Disallow: / ```

The GPTBot group expresses the training-use restriction described by OpenAI. The Google-Extended group also restricts its documented Gemini grounding use, so this example cannot be described as blocking only training across both providers. Neither group names Googlebot, Bingbot, OAI-SearchBot or PerplexityBot; review their existing access rules separately before changing the file.

OpenAI explicitly documents independent settings for OAI-SearchBot and GPTBot. This gives a documented way to retain ChatGPT search access while expressing a restriction on model-training use. It does not establish that past training has been reversed or promise future answer placement. Perplexity describes PerplexityBot as a search crawler rather than a foundation-model training crawler; blocking it would target the documented search use.

03

User requests need a separate policy decision

Do not treat a search-crawler rule as a complete control over a user's requested page fetch. Use the Google, OpenAI and Perplexity user-triggered entries above when deciding which requested fetches the site intends to support. OpenAI says robots.txt rules may not apply to ChatGPT-User actions. Perplexity's wording is similarly qualified:

“Since a user requested the fetch, this fetcher generally ignores robots.txt rules.”

Perplexity crawler documentation

Keep those qualifications in the policy record. They do not mean every user fetch will succeed, and they do not turn these user-triggered fetchers into a training-control token. The distinction between access and broader answer visibility sits within the GEO overview.

04

Check the served policy and the access layers

After a change, retain the actual publicly served robots.txt response, its retrieval time and the text containing the intended groups. Compare that response with the approved edit; an edited local file alone is not evidence of what a crawler can retrieve. Record the intended action for each token and check that existing rules do not undermine it.

Also review CDN and WAF restrictions against the intended access. Perplexity notes that sites using a WAF may need to explicitly permit its bots. Keep the resulting access-check evidence alongside the file response. The existing ChatGPT access checklist covers the OAI-SearchBot check; use that checklist for this layer. Access-policy changes are one possible condition to examine in a Perplexity visibility investigation, without treating them as proof of the cause.

Allow for the delay the provider actually documents: OpenAI says search systems can take approximately 24 hours to adjust after a robots.txt update; Perplexity says changes may take up to 24 hours to be reflected. These are qualified adjustment windows, not deadlines for appearing in answers. Keep the provider's wording and the edit time beside the follow-up result instead of assigning one universal delay to all four services.

This article was drafted with AI assistance and reviewed by our team before publishing.

05

Want to see how AI engines describe your brand today?

Get a free growth audit