Can You Block AI Training and Still Rank on Google?

Writer: Ryan Frankel

Editor: Lillian Castro

Reviewer: Jordan Sprogis

Follow the HostingAdvice team for a daily dose of tech news, trending IT discussions, and interviews with the web's most innovative technologists.
Follow Us:
2.7k
1k

A week ago, a post made it to the top of r/Cloudflare with a title guaranteed to make a lot of site owners spit out their coffee: “PSA: CloudFlare Now Defaults To Allowing AI Scraping Of Your Sites.” The poster went into their Cloudflare dashboard and found an odd toggle. In the past, there was a single “Block AI Bots” toggle here, but now the Training setting showed “Allow.”

On September 15, Cloudflare put an end to what it calls the search-or-AI-training tradeoff. Here’s the upshot: Now you can ask Applebot, Bingbot, and Googlebot to index your site for search, but not to use it in training. And, as Apple, Google, and Microsoft have publicly said, they won’t penalize you in search if you turn them down for training.

Why One Switch Was Never Going to Work

Cloudflare calls bots that do both search and training “mixed-use” crawlers. And they include, for example, Googlebot, Bingbot, and Applebot. They now account for 36.6% of all verified crawler traffic on the network, the single largest category Cloudflare tracks.

1% of sites on Cloudflare block search bots

And here’s the thing: Site owners tend to have very different views of each part of the bot. Fewer than 1% of Cloudflare’s sites block search bots, but 17% have some mechanism in place to prevent training. In the past, the “Block AI Bots” toggle would flip off both parts of the bot. And so anyone using it on, say, an ad-supported site would be risking their Google traffic in hopes of keeping their content out of some AI training run.

“Refuse one, and you refuse the other,” the technical blog post that went along with this change said.

Now there’s a Training setting in Cloudflare called Disallow AI Training, named for the Disallow directive it publishes in your robots.txt. And in effect, it means Cloudflare’s network publishes the no-training preference, and Accountable mixed-use crawlers keep crawling your site for search purposes. Every other training crawler is blocked by Cloudflare. That includes the training-only crawlers of Amazon, Anthropic, Meta, and OpenAI. And none of those crawlers touch search.

What “Accountable” Actually Requires

In this context, being “accountable” is what it means when a bot/operator is transparent and gives website owners meaningful control over how their content used by AI systems. An operator would need (or need to commit to over a stated time frame) four things: a way to opt out of training (through robots.txt or a similar standard), a way to opt out of AI-generated summaries of its content, visibility at the URL level of which pages were made available for training, and a statement that opting out of training won’t impact traditional search results.

How Cloudflare got from 'Block AI Bots' to Accountable crawlers

July 1, 2025
First Content Independence Day: Cloudflare flips new domains to block AI crawlers by default and launches Pay Per Crawl
July 1, 2026
Second Content Independence Day: Search, Agent, and Training become separate controls, with new ad-supported defaults set for Sept. 15
Sep. 15, 2026
Accountable designation and Disallow AI Training launch; Block AI Bots and Managed Robots.txt are deprecated in favor of the three controls and Bot Preference Sync
Early 2027
Microsoft targets honoring a domain-level no-training robots.txt preference for Bingbot; Cloudflare aims to add one-place control over how much content appears in AI summaries

Apple, Google, and Microsoft now meet that standard (though with some caveats). And so do Amazon, Anthropic, Meta, and OpenAI. Their training crawlers are already separate from their search crawlers. So Cloudflare can block one without affecting the other.

You can find the list of Accountable operators on Cloudflare Radar. And there, the company says it will keep track of whether they live up to their commitments.

In the announcement, Cloudflare CEO and co-founder Matthew Prince framed it more broadly: “This is how we make the Internet better: preserving the openness that makes search valuable while giving the people and businesses behind the web meaningful control over how their work is used,” he said.

Now, About Those Caveats

First, “Block” now means more than it did two weeks ago. Until now, Block and “Block on pages with ads” would skip over mixed-use crawlers. That’s because doing otherwise would risk affecting search. Now they apply to all training crawlers, including mixed-use ones. If you click Block on the Training setting in your dashboard, Bingbot, Googlebot, and Applebot no longer visit your site. Not even for search. If you’d like to turn off training but keep search, the only option is Disallow AI Training.

Second, Bing isn’t quite there. Microsoft is building support for a no-training preference in robots.txt at the domain level. That’s targeted for early 2027. Until then, Disallow AI Training doesn’t automatically pass this preference along to Bingbot. In that case, site owners would need to use Bing’s content-removal tool or its NOARCHIVE meta tag. Apple doesn’t yet have URL-level inspection. According to Cloudflare, Apple has shared an in-progress solution for next year.

Third, the defaults. They’ll automatically carry forward from your current settings. And in almost every case, Cloudflare says customers won’t need to change anything. If you currently have Block in your settings, it will become Search: Allow, Training: Disallow AI Training, and Agent: Block on pages with ads. For a new site, Cloudflare asks if you sell ads on your pages. If you do, it sets training to Disallow AI Training and agents to Block on pages with ads. If not, everything stays on Allow.

And that brings me to that Reddit post. “That’s strange, because their blog has the recommended, default setting as ‘Disallow’ and their email also states ‘most customers won’t need to change anything,'” the user wrote. Both statements are true; they just describe two different presets, and the “I monetize pages that serve ads” checkbox decides which one you get.

Training Was the Hard Part. Summaries Are Next.

Cloudflare’s blog post says mixed-use crawlers were “the hard part of the training question.” But right after that it moves on to the next question: AI summaries. And one of the four requirements for Accountable is an opt-out of summaries. Cloudflare says a site-wide yes/no answer is too coarse, and its aim in pushing this is to let each site decide, once on Cloudflare, how much of its content appears in an AI summary, rather than with each operator separately, by early next year.

There’s good reason for that push. More than half of consumers read AI summaries in search, and they’re also more than 40% more likely to end their search without clicking through. That’s not good news for sites that get paid per page view, but visitors sent by AI search convert 3x-5x the rate of those sent by traditional search. Fewer visits, sure, but more intent. Cloudflare’s point is that it isn’t going to pick one setting for everyone. A site aiming for reach would probably set a different one from one hoping to sell.

But underlying this push is a promise from three of the world’s biggest crawler operators: They have publicly said your rankings won’t change if you add a no-training flag to your site. Right now, the only way to check if they honor that promise is through Cloudflare’s scoreboard.

About the Author

CTO and Contributing Expert

Ryan Frankel has been a professional in the tech industry for more than 20 years and has been developing websites for more than 25. With a master's degree in electrical and computer engineering from the University of Florida, he has a fundamental understanding of hardware systems and the software that runs them. Ryan now sits as the CTO of Digital Brands Inc. and manages all of the server infrastructure of their websites, as well as their development team. In addition, Ryan has a passion for guitars, good coffee, and puppies.

« BACK TO: BLOG

Meet the Experts

Our team of experts with a combined 50+ years of experience in web hosting serve insight and advice to more than 20 million users!

We Know Hosting

$

4

8

,

2

8

3

spent annually on web hosting!