Cloudflare now lets you block AI training on your website — without disappearing from AI answers
Cloudflare's new AI Crawl Control splits one blunt setting into three switches: search, AI training, and AI agents. What the Sept 15 update changes for anyone publishing content.
This morning an email from Cloudflare landed in my inbox that most site owners will scroll past. They shouldn't. As of September 15, Cloudflare gives every website behind its network a separate switch for AI training — you can now stay visible in search and AI answers while telling training crawlers to stay out.
I recorded a short video on what changed and what it means in practice:
One setting became three switches
Until now, "block AI bots" was a sledgehammer. If you told Cloudflare to keep AI crawlers out, you kept all of them out — including the ones that decide whether your site shows up when someone asks ChatGPT or Claude a question. So most of us left the door open, because losing AI discoverability is not a trade anyone in marketing wants to make.
Cloudflare's update splits that one decision into three: search, AI training, and AI agents. The new "Disallow AI Training" setting publishes the preference in your robots.txt and blocks the training-only crawlers run by the big model companies, while accountable mixed-use crawlers remain allowed for search. Your content keeps getting found; it stops being raw material.
Why this matters more coming from Cloudflare than from anyone else: roughly a fifth of all web traffic passes through their network. They are the bouncer at the door for a very large part of the internet. When the bouncer changes the house rules, the rules actually change — no lawyers, no license negotiations, just a setting in a dashboard.
The part everyone will get wrong
Blocking AI training does not mean ChatGPT or Claude can no longer open your site when a user asks about it. Agent access and search access are separate switches. You can be fully present in AI answers, assistants, and agent workflows while still saying: don't bake my work into your next model. That distinction — training versus access — is the whole point of the update, and it's the part I expect most people to miss.
Ethan Mollick put his finger on the underlying issue in Co-Intelligence:
"Even if pretraining is legal, it may not be ethical. Most AI companies do not ask for the permission of the people whose data they train on."
Nobody asked. Now, at least for sites behind Cloudflare, you can answer anyway.
The judgement call this creates
Here's where it gets interesting for operators. This switch forces a question most content teams have never had to answer: which part of your content is a moat, and which part is distribution?
If your site regurgitates what's already on the internet, blocking training protects nothing — there's no original signal in there to protect. If your site carries real operator knowledge — pricing logic, process detail, hard-won field notes — you now get to decide whether that becomes free training material. It's the same editorial discipline I ran into when I shipped a magazine section and spent most of the day deciding what not to index: the work is rarely the publishing, it's the deciding.
My starting position for client sites: leave search on, decide training case by case, and watch the agents switch closely — because agent traffic is about to become a real acquisition channel, the same way platforms are already opening their ad tooling to AI assistants. Blocking agents today feels safe and will look like turning away customers tomorrow.
One more practical note: this is infrastructure, so treat it like infrastructure. A setting you flip once and never audit is how measurement quietly breaks — I've written about what happens when automated routines run without anyone checking the context they run in. Put the crawl settings on the same review cadence as your tags and your consent banner.
The tools keep getting better at doing. The scarce skill is still knowing which switch to flip.
Sources & further reading
External
Cloudflare: Have it both ways — stay discoverable in search while disallowing AI training (the announcement, Sept 15)
Cloudflare Docs: AI Crawl Control
Ethan Mollick, Co-Intelligence (Portfolio, 2024)
Related posts
I shipped a magazine section in a day. The work was deciding what not to index.
Meta just opened Ads Manager to Claude — here's the setup that actually works
My scheduled Claude Code routine failed three weeks running