Cloudflare changed how it handles automated crawlers just under a week ago, and new sites joining the network now inherit stricter AI settings without lifting a finger.
Let's start with what Cloudflare is: a company whose infrastructure sits in front of a large share of the world's websites, filtering the traffic that reaches them.
Put simply, its default is protocol across a large slice of the web.
What actually changed
The network now sorts automated traffic into three buckets: Search, Training and Agent.
For pages that carry advertising, it permits Search while blocking Training and Agent by default on new domains.
The distinction is important. Search bots index a page so people can find it, sending readers back; Training bots ingest the page to build AI models and send nothing back; Agent bots act on a user's behalf, often bypassing the site entirely.
Cloudflare put the rule plainly in its developer documentation, saying bots classed as Training or Agent are blocked on pages that display ads while Search stays allowed.
The company then re-evaluated every page under the new rules the following day, applying the classification across the network.
Why the default is the story
The significance lies less in the block than in the direction of the default, and who now has to act.
Until now, the burden fell on publishers to notice AI crawlers, understand them and write rules to keep them out, an effort most never made, which let training bots harvest content for free by inaction.
Flipping the default reverses that. Silence now means exclusion rather than consent, and it is the AI companies, not the publishers, who must seek access.
That is a meaningful shift in leverage over who pays whom for web content, and it scales automatically to every new site that joins.
The commercial logic underneath
The defaults do not stand alone. They pair with tools that turn the block into a toll.
Cloudflare shipped management controls, including Bot Preference Sync and a site-level Disallow AI Training setting, alongside monetisation features such as Pay Per Crawl and "Cloudflare Wallets" that let publishers charge automated agents for access.
Read together, the sequence is deliberate: build the wall by default, then sell the gate. A blocked crawler is a crawler that can be charged, and Cloudflare positions itself as the tollbooth in between.
The limits worth noting
Two caveats temper the reach.
The change is not retroactive in force: Cloudflare signalled the policy on 1 July and did not reset sites that already carried custom rules, so existing domains keep their settings unless operators change them.
And an independent audit tracked in the rollout followed 20 named AI bots against the new taxonomy, a reminder that the system is only as good as its classification, and that a bot's bucket, not its behaviour, decides what it may take.