When AI Control Becomes an SEO Decision
Cloudflare wants to give publishers more control over AI crawlers. But when Googlebot is classified as both Search and Training, a choice aimed at AI training can also become a question of search access, visibility and traffic.
NextNet’s earlier Cloudflare investigation examined the company’s broader AI-web strategy. This follow-up asks a narrower practical question: what happens when a publisher’s decision to limit AI training also touches a crawler used for search?
Cloudflare classifies Googlebot as both Search and Training. Google, meanwhile, describes Googlebot as the crawler used for Google Search and offers Google-Extended as a separate robots.txt product token for specified AI-related uses. That is a documented difference in how crawler use is described. It is not proof of a corporate conflict, incorrect information, or an automatic loss of ranking.
When AI control becomes an SEO decision
Cloudflare lets customers set policies for Search, Agent and Training behaviour. For publishers, that can look like a way to restrict model-training use without excluding all automated access. But Cloudflare’s documentation also says a crawler can have multiple purposes, making its behaviour classification consequential.
What changes on September 15?
Cloudflare says that, from September 15, 2026, new domains will by default block bots classified as Training or Agent on pages where ads are detected, while Search remains allowed. Customers can choose other settings. Cloudflare also says customers may opt out before the new defaults take effect. For existing free customers, the actual configuration should be checked in Cloudflare’s dashboard, because the public material does not establish the final settings for every account.
Search, Agent and Training
Cloudflare describes Search as crawlers collecting or indexing content to answer questions later; Agent as automated activity acting in real time on a person’s behalf; and Training as crawlers collecting content to train or fine-tune a model. Its documented mitigation options are block on all pages, block on pages with ads, or allow.
Cloudflare classifies Googlebot as Search and Training
Cloudflare states that mixed-purpose crawlers combining Search and Training are blocked by configurations that block AI training. Its documentation names Googlebot as an example. Cloudflare describes this as applying rules based on all of a crawler’s behaviours. That does not by itself establish that Google Search will lose access to any particular page, but it makes the classification material when a publisher selects a restrictive policy.
Google describes Googlebot as the crawler for Google Search
Google describes Googlebot as the crawler that keeps Google Search results fresh and up to date. Google says blocking Googlebot can affect Google Search, including Discover, Images, Video and News. Google also distinguishes crawling from indexing: crawling discovers and fetches content; indexing is subsequent processing. Neither indexing nor ranking is guaranteed for a particular page.
What is Google-Extended?
Google-Extended is, according to Google, a standalone robots.txt product token rather than a separate HTTP user-agent. It lets publishers manage whether content Google crawls may be used for training future Gemini models and certain grounding uses. Google says Google-Extended does not affect inclusion in, or ranking in, Google Search.
Two descriptions of technical access
Cloudflare describes the uses a crawler may have and applies policies to those behaviours. Google describes how publishers can manage uses within Google’s own products. For a publisher, the difference can still have practical implications for crawler access, visibility and traffic. This is a possible technical and strategic tension, not evidence that either company’s description is false.
Crawling, indexing and ranking are different
Blocking can affect a crawler’s ability to fetch content, but it does not independently determine whether a URL is indexed, shown or ranked. Google describes Search as a multi-stage process and says ranking is programmatic and uses many signals. No automatic ranking or traffic loss should be inferred from a Cloudflare setting.
Cloudflare’s infrastructure role
NextNet uses the term “control layer” as an analytical description, not as a Cloudflare product name. When an infrastructure provider both classifies traffic and provides the policy controls, its definitions matter to the customer’s decision. The prior article covered payment models in detail; this article focuses on the transparency required before a classification may affect access.
The Cloudflare–OpenAI research pilot
On July 8, 2026, Cloudflare and OpenAI said they had launched a research pilot to explore whether network signals from participating websites could help AI search engines discover and index relevant open-web content. Cloudflare named content freshness, traffic quality and actual page changes. This is the companies’ description of a pilot. Public material does not fully explain selection, the exact signals, data handling, equivalent access for other providers, or any production impact.
What publishers should check before September 15
- Active Cloudflare policies for Search, Agent and Training.
- Whether ad detection is used and which pages it covers.
- How mixed-purpose crawlers, including Googlebot, are shown and treated in the account.
- robots.txt rules for Googlebot and Google-Extended as distinct controls.
- Search Console and analytics baselines before any change, without treating variation alone as proof of causation.
The Questions NextNet Sent – and That Remain Unanswered
NextNet sent detailed questions to Cloudflare and OpenAI on July 31, 2026. NextNet emailed Cloudflare at 03:15 and OpenAI at 03:17. Both companies were asked to respond by August 4 at 17:00 CEST and were invited to request additional time. The questions were not answered within the stated deadline. NextNet is therefore publishing every question in full so readers can assess for themselves which issues remain unresolved.
The purpose is not to draw conclusions about the companies’ motives, but to allow readers and publishers to assess for themselves what remains unclear.
Questions sent to Cloudflare
- How does Cloudflare determine whether a crawler is classified as Search, Agent or Training, particularly when the same crawler may serve multiple purposes?
- Is there a public process for crawler operators or website owners to challenge, correct or appeal a classification? If so, what is the process and expected response time?
- How are participating websites selected for the OpenAI research pilot, and does each website owner actively opt in?
- Which exact signals are provided to OpenAI through the pilot? Do they include URLs, content freshness, page changes, traffic-quality metrics, raw traffic data, IP addresses, visitor information or user queries?
- Will other AI search and answer-engine providers be offered access to the same signals, technical interfaces and commercial terms as OpenAI? What criteria must a company meet to participate?
- Can payment, a commercial agreement or participation in a pilot affect crawling priority, indexing, source selection, citation frequency or placement in AI-generated search results?
Additional question: How does Cloudflare distinguish between paid access to content and any possible influence on ranking or source prioritisation?
Questions sent to OpenAI
- Which exact signals does OpenAI receive from Cloudflare through the pilot? Do these include content freshness, page changes, traffic-quality information, URLs, traffic data or other metadata?
- Cloudflare states that OpenAI contributes real user queries for testing. How are these queries used, and are they shared directly with Cloudflare? Are they aggregated, de-identified or anonymised?
- Is any information from the pilot used for model training or fine-tuning, or is it limited strictly to search, crawling and indexing research?
- Is the pilot currently affecting ChatGPT Search in production, including crawling priority, indexing, source selection, citation or result placement?
- Can payment to a publisher, a commercial agreement or participation in a partner programme influence whether a source is crawled, indexed, cited or prioritised?
- How does OpenAI distinguish between pay-to-access and pay-to-rank? Can you confirm whether payment for crawler access provides any ranking, citation or placement advantage?
- Is OpenAI’s participation exclusive, or does OpenAI support making the same Cloudflare signals and technical interfaces available to competing AI search and answer-engine providers on equivalent terms?
- If a publisher accidentally blocks OAI-SearchBot and later removes the block, how quickly can the content become eligible again for discovery, summaries, snippets and citations?
Conclusion: control requires transparency about consequences
Cloudflare’s rules may give publishers more agency. But when infrastructure and a crawler operator describe use in different ways, settings need to be examined as more than a simple yes-or-no choice about AI. What is publicly missing is the level of visibility that would make the consequences predictable.
💬 What do you think?
What transparency should a publisher be able to require before a crawler classification can affect access and visibility?
Share your thoughts in the comments.
📚 Related articles
Vad är din reaktion?
Gilla
0
Ogilla
0
Kärlek
0
Rolig
0
Wow
0
Ledsen
0
Arg
0
Kommentarer (0)