Cloudflare has introduced a Disallow AI Training setting that lets a site refuse AI training crawlers while remaining indexed by search engines, a split that was not possible before. The change took effect on September 15. Apple, Google and Microsoft either honor the setting or have committed to honor it within a specified time frame.
The problem it addresses is that some crawlers serve two purposes at once: they collect pages for search results and also feed material into AI training. Until now, turning away one use meant turning away the other, so a site that wanted out of AI training risked dropping out of search. The new control separates the two for Applebot, Bingbot and Googlebot.
Sites that want to keep AI training crawlers out can still use Block, but that now also stops search crawling. Disallow AI Training instead keeps Accountable mixed-use crawlers available for search while blocking all other training crawlers, including the training-only bots of OpenAI, Meta, Anthropic and Amazon, which Cloudflare says can be shut out without any search impact. The setting is available only for the Training category, not for Search or Agent.
Cloudflare is retiring the older Block AI Bots control, and customers of Managed Robots.txt are being moved onto the new system. Existing preferences carry over automatically, and the company says almost no one needs to do anything. The legacy Block AI Bots setting maps onto the new, more granular controls.
New domains were offered two presets starting September 15, and which one a site gets depends on whether it runs ads. Ad-supported sites receive the more restrictive option. The choice can be changed during onboarding or afterward, and the controls apply at the domain level.
The system sorts crawler behavior into three categories: Search, Training and Agent. Agent covers user-directed chat fetching and browser-use bots, and there is no Disallow option for agents yet because no well-established directive for agent preferences exists. Disallow AI Training works by publishing a Disallow directive in robots.txt.
An ads-only preference cannot be expressed in robots.txt, so Cloudflare detects ad-serving pages itself, though the list is too large to enumerate. The company says robots.txt alone cannot identify crawlers or stop those that ignore it, so it publishes the preference, classifies crawlers and blocks the ones that ignore the signal. Operator behavior is reported on Radar.
Cloudflare also introduced an Accountable designation for operators that meet certain conditions. It requires training opt-out through robots.txt or a similar standard, an opt-out for AI summaries with the operator and later through Cloudflare, URL-level visibility into training pages and search metrics, and assurance that opting out will not hurt traditional search. Apple, Google and Microsoft meet the qualifications. Mixed-use crawler operators already have to offer an AI summaries opt-out, and Cloudflare's goal by early next year is to let sites control how much of their content shows up in summaries.
The announcement follows Cloudflare's BotBase directory and Business Insights, both unveiled on its second Content Independence Day, and the Bot Preference Sync feature introduced on August 21. Cloudflare says 17% of the sites it serves have switched on some kind of training block, while fewer than 1% block search bots. Talks with operators began in July.













