Which AI crawlers should you allow? The defaults are changing

Cloudflare blocks AI training crawlers by default from 15 September. What the three crawler types do, what blocking each one costs, and who decides.

Blog Collection Athour img
Michael Forystek
Co-founder, Growth & Partnerships
shape

On 15 September, Cloudflare changes what happens to AI crawlers arriving at a large share of the web. Training and agent crawlers get blocked by default; search crawlers stay allowed. The change applies to new Cloudflare customers, new sites set up by existing customers, and all existing free accounts, on pages that display advertising. A configured paid zone sees nothing happen at all.

For most organisations reading this, the direct effect on the fifteenth is probably nothing. That is worth saying plainly, because the interesting part is not the blast radius. It is that the decision about who may read your website has quietly moved from a text file you write to a default someone else sets — and that the default has flipped from yes to no.

Three jobs, not one

Until recently, crawler policy was binary: a bot was either allowed or it wasn't. That framing has broken down, because the companies sending the bots now send three different kinds.

Training crawlers collect content to train future models. GPTBot, ClaudeBot, Google-Extended, CCBot and their peers. Blocking them keeps your material out of the next model generation and removes you from nothing a user sees today.

Search-indexing crawlers build the index an assistant searches when answering a question — OAI-SearchBot, Claude-SearchBot, PerplexityBot. Blocking these is the consequential move: it takes the organisation out of the pool of sources an AI assistant can cite.

Agents and user-directed fetchers retrieve a page because a person asked for it in that moment: ChatGPT-User, Claude-User, Perplexity-User. Someone pastes your URL into an assistant and asks what it says. Blocking these refuses a request a real prospect made about you.

OpenAI and Anthropic each name their three bots separately and document the split. Cloudflare's new controls use the same three categories. When the model providers and the infrastructure layer independently arrive at the same taxonomy, it has stopped being a technical curiosity and become the shape of the decision.

What actually changes on the fifteenth

The precise scope matters more than the headline, so here it is without embellishment. Search crawlers remain allowed under the new default — the category that determines whether you can be cited in an AI answer is the one Cloudflare left open. Training and agent traffic is what gets blocked, and only on ad-showing pages, and only for accounts that never configured the setting.

A regulated institution running a corporate site on a paid plan with no advertising is, on this reading, untouched. Anyone can opt out through the security settings before the date.

So the honest summary is not that a wave of sites goes dark. It is that the largest infrastructure provider on the web has decided the sensible default is to charge for training access and give away search access — and that publishers who never thought about the question now have a position by inheritance rather than by choice.

What blocking actually costs

The decision deserves more than a default in either direction, because both sides carry weight.

The case for declining training access is straightforward. Content is an asset the organisation paid to produce, and a model trained on it generates derivative value the organisation never sees. For firms whose published material is a genuine competitive artefact — proprietary research, original methodology, analysis nobody else has done — declining costs almost nothing in visibility, because the search category is separately controllable.

The case for staying open on retrieval has grown considerably stronger. When a compliance officer asks an assistant which providers handle a particular problem, the answer draws on sources the assistant can reach. Absence from that index is absence from the shortlist, and it is invisible: there is no ranking report that tells you it happened. For an organisation whose buyers research this way — and in financial services, increasingly they do — being reachable is the whole game.

The two are no longer in tension, which is the genuinely useful development. An organisation can decline to feed model training and remain fully citable in AI answers. Two years ago that choice could not be expressed.

The caveat nobody advertises

Robots.txt is a convention, not an enforcement mechanism. Providers often treat user-directed fetches as access on behalf of a person rather than crawling, and apply crawl rules differently as a result. Reputable operators honour the file; the market's less reputable end is documented as honouring it selectively.

Two consequences follow. Robots.txt states a position rather than building a wall, which is precisely why the control point is migrating to the network layer, where blocking is enforced rather than requested. And an organisation with genuinely confidential material should not be relying on a text file to protect it: that is an access-control question, and answering it with robots.txt is a category error.

The decision, as four questions

Institutions that handle this deliberately settle the same four questions in writing:

  • Who owns it? The decision sits between marketing, legal and IT, which in practice often means it sits nowhere.
  • Where is it actually made? Increasingly not in robots.txt but in the CDN or security layer, where nobody thinks to look — and where, from this month, a default may already have answered on your behalf.
  • What is the position on training versus retrieval? These are separable now, and a considered answer usually differs between them.
  • How would you know it was working? Retrieval presence is measurable — assistant citations, referral traffic from AI tools, natural-language branded queries appearing in search data — and almost nobody measures it. The measurement is the part most organisations skip, which is why most positions on this are held without evidence.

If your institution has not looked at where this decision is currently being made, that review is a short conversation.

Where this lands

Nothing dramatic happens to most websites on 15 September. What happened is that a default changed, in a place most organisations do not look, on a question most have never formally answered — and defaults, unlike policies, apply whether or not anyone agreed to them.

The organisations that will be found by AI assistants in a year are not the ones that blocked everything or allowed everything. They are the ones that noticed the question had three parts, took a position on each, and know where that position is written down. Everyone else has a policy too. Theirs is whatever their infrastructure provider decided this month.

Related reading:

Sources: Cloudflare changelog and blog, "New options to manage AI traffic" (1 July 2026); OpenAI and Anthropic published crawler documentation. Scope as announced at the time of writing.

Ready to Own Your AI?

Stop renting generic models. Start building specialized AI that runs on your infrastructure, knows your business, and stays under your control.