AI Business

Cloudflare Launches Tool to Block AI Training Bots While Preserving Search Access

Website owners can now use Cloudflare's new Disallow AI Training setting to prevent AI companies from scraping their content for model training while still allowing search engines to crawl normally.

·2 min read
Cloudflare Just Gave AI Training Bots the Middle Finger
Cloudflare Just Gave AI Training Bots the Middle Finger

Most site operators welcome search engine crawlers like Googlebot, yet many wish to prevent AI firms from harvesting their material for training purposes. Historically, achieving this distinction has proven difficult. Cloudflare is now addressing this challenge head-on.

On September 15, Cloudflare rolled out a Disallow AI Training feature that permits website proprietors to maintain access for conventional search crawlers while simultaneously rejecting AI training crawlers. This represents more than a gentle suggestion to bots.

The infrastructure company reports that training-specific crawlers operated by Amazon, Anthropic, Meta and OpenAI can be completely blocked, whereas dual-purpose crawlers including Googlebot retain the ability to index sites for search purposes. Google has already provided publishers with mechanisms to restrict certain Gemini training and grounding applications through Google-Extended without impacting their visibility in Google Search or search rankings.

Cloudflare is extending its approach further for websites that rely on advertising revenue. Its default configuration now essentially instructs: permit search engines, reject AI training crawlers, and reject AI agents on pages where advertisements appear.

The system does have constraints. Since robots.txt depends on bots complying voluntarily, enforcement cannot be guaranteed. Additionally, blocking Google-Extended will not prevent your material from appearing in Google's AI Overviews or AI Mode, which are classified as components of Google Search itself.

The underlying principle, however, is unmistakable. For decades, website operators permitted bots to access their material because they received something tangible in exchange: referral traffic. Artificial intelligence threatens to alter this arrangement by extracting content without necessarily directing visitors back to the source.

Cloudflare is offering publishers a straightforward path forward: maintain the search engine traffic while firmly rejecting AI training bots.