The Cloudflare box you ticked in 2024 starts blocking Googlebot on 15 September

by•

The legacy "Block AI bots" switch is deprecated on 15 September, and mixed-purpose crawlers get folded into it. Googlebot, Applebot and BingBot are in scope. Opting out is one click and the window is 25 days.

Sometime in the last two years a lot of us flipped a switch in Cloudflare called "Block AI bots." It was one click, it was free, it was framed as protecting your content, and there was no reason not to. I did it. I never thought about it again.

On 15 September that switch is deprecated, and the thing it does changes.

Here is the specific mechanic, from Cloudflare's own docs rather than the coverage. Cloudflare now sorts automated traffic into three behaviours: Search (indexing your content to answer questions about it later), Agent (something acting in real time on a person's behalf, a chat fetch bot or Claude driving a browser), and Training (absorbing your content into a model). Until now, the "Block AI bots" preset deliberately excluded crawlers that do more than one of those jobs. That exclusion is what kept Googlebot out of scope, because Googlebot crawls for search and for training with the same bot.

On 15 September that exclusion goes away. Mixed-purpose crawlers that combine Search and Training get blocked by every configuration that blocks AI training, including the legacy preset. Cloudflare names the three you'd care about: Googlebot, Applebot, BingBot.

So if you turned that on and forgot, in 25 days you are telling Google not to crawl you. Google is still roughly 88% of referral traffic on Cloudflare's own numbers. That's not a content-licensing decision, it's an accident.

The opt-out exists and takes about a minute: Security Settings, then Configure AI bot policies. You can confirm you want no change to Training crawlers that also crawl for Search, and you can set each of the three categories independently, with three levels each: block everywhere, block only on pages that display ads, or allow.

Two things I want to be accurate about, because the hot takes I've read are not.

First, the widely-repeated "Cloudflare is blocking AI agents by default" is narrower than it sounds. The new default, Training and Agent blocked with Search allowed, applies to pages that display ads, and to new domains onboarding. If you run a normal product site with no ad units, that default barely touches you. The change that actually reaches into an existing account is the mixed-purpose reclassification above, and it reaches you only if you have training-blocking switched on. Which, if you're the kind of person who ticks the safe-looking box, you do.

Second, Cloudflare sells bot management and is building a market for paid crawling. Every number in their report is measured by a company with a position in the outcome. Read them as direction, not measurement. The direction: more than half of internet traffic is now non-human, 52% of crawler requests were for AI training as of June 2026 against 22% in spring 2025, and mixed-use crawlers are over 36% of activity. That last one is the whole story. One bot, several jobs, and the taxonomy everyone's blocking rules were written against no longer describes it.

The part I think is genuinely strategic, and it isn't the Googlebot thing.

Agent is now a category you can block, which means it's a category you have to have an opinion about. A browser-use agent hitting your pricing page is not a crawler harvesting you. It's a person who asked something to go look at your product, sitting there waiting for an answer. Blocking Training is a stance about your content. Blocking Agent is a stance about a customer. Those got put on the same settings page next to each other, in the same visual language, and they are not the same decision at all.

I don't have a satisfying answer for Murror. We're not ad-supported, so the default doesn't bite. But the honest reason I'd never block Agent is thin: I want to be findable, and right now agents are how a growing share of finding happens. If people started pointing agents at private journalling products expecting to get in, I'd feel very differently, very fast. I'd rather say that than pretend I've reasoned it through.

What I'd actually do this week: open Security Settings, look at what's on, decide each of the three categories on purpose, and write down why. Not because Cloudflare's defaults are wrong, mostly they're reasonable, but because a default you never chose becomes a strategy by attrition. This one has a date on it.

107 views

Add a comment

Replies

Best

I run marketing for a handful of app sites behind Cloudflare and nobody on my side remembers which ones got that switch flipped in 2024, which is the real problem here. Did Search Console crawl stats give you any warning, or is checking the header the only reliable way to know before the 15th?

  Search Console gave me nothing, and I want to be precise about why, because it is the part that cost me the most time.

Crawl stats will show you a drop in requests, but a drop has many parents. Fewer requests looks identical whether you were blocked, deprioritised, or your own server started erroring. In my case it was the third one and I read it as the second for weeks. So no, it did not warn me, and I would not rely on it to warn you either.

Checking directly is the reliable way, but I would not check only the header. Two things, both cheap:

First, curl your own site with Googlebot's user agent, and do it against your ugliest URLs rather than the homepage. The homepage is the page everyone tests and the page that is always fine. Mine broke only on paths containing a dot.

Second, and this is the one that actually settles it, use the URL Inspection tool in Search Console and hit Test Live URL. That fetches as Google, right now, and it tells you what Google receives rather than what you think you are serving. If something upstream is intercepting, that is where it shows up.

For finding which sites had the switch flipped, there is no clever route I know of. Someone has to open each dashboard and look. The awkward part is that the setting was a default nobody chose, so nobody remembers choosing it, and that is exactly why it survived two years without review.

  Your point that a drop has many parents is the bit I'll be repeating to the developers I work with, because I would have read it as a block too. Thanks for being that precise about it.

 Serdar's answer is the right one for diagnosing after the fact, and I'd add the reason it can't help you before the 15th: nothing has happened yet. The reclassification isn't live, so there's no blocked request for any crawl report to show you. Every diagnostic in this thread detects a block that already cost you something. For a deadline that hasn't arrived, reading the config is the only check that exists.

On "nobody remembers which ones" — that's the actual problem and it's not really a Cloudflare problem. Serdar is right that someone has to open each dashboard. One thing that may save you the clicking if the sites sit under one account: Cloudflare's API can read zone settings per zone, so a token plus a loop over your zone list can tell you which ones have bot-blocking on without anyone logging in. I haven't run it against this specific setting so I don't want to promise the field is exposed the same way for every plan — worth ten minutes with a read-only token before you commit an afternoon to clicking.

And I'd write the answer down somewhere the next person finds, per site, because the thing that made this bite is that it was a decision nobody recorded making.

  That's the distinction I was missing, since I'd been treating the 15th as something to watch for rather than something I have to go and read. The API loop is what I'll try, thanks. Those zones were set up over years by different people, so how would the loop catch a domain where someone wrote a custom WAF rule instead of flipping the toggle?

  It wouldn't, and that's a fair hole in what I suggested. Zone settings and custom rules are separate objects, so a loop that reads the toggle comes back clean on a zone where someone hand-wrote a rule doing the same job.

You can extend it — list each zone's rulesets and grep the expressions for bot fields and user-agent matches — and that catches the ones written the obvious way. It won't catch a rule that blocks the same traffic without naming it: by ASN, by a UA substring, by a rate limit a crawler trips and a person doesn't. "Blocks Googlebot" is a behaviour, not a field, so no grep is complete.

For the domains that actually matter I'd test the outcome instead of the config. One caveat there, because I'd have got this wrong myself: curling with Googlebot's user agent doesn't tell you what Googlebot gets. Cloudflare verifies bots by reverse DNS, so your curl is an unverified bot pretending, and rules can treat those two very differently in either direction. The fetch that isn't a guess is Serdar's — URL Inspection, Test Live URL, which goes out as the real thing. Config loop to triage everything, live fetch on the handful you'd be upset to lose.

Founder here, I run usebravery.com, so I have been staring at exactly this problem for a week and I am not neutral about crawl access.

Adding a site-owner-side version of Omri's point about invisible failure, because the Googlebot case is less self-announcing than it looks.

The assumption in "at least your search traffic actually drops and you notice" is that you had search traffic to lose. A lot of the sites that ticked this box are younger than that. On ours, Search Console showed 230 URLs sitting in Discovered, currently not indexed, which is Google saying it knows the URL exists and has not fetched it. That state looks the same whether the cause is a crawl block, a crawl budget decision, or your own server handing back errors. There is no line in any report that says something upstream refused me. You get the same grey row for all three.

We went through that this week. The AI bot toggle was one of the things we turned off, and it was the right call regardless. But it was not our cause. Our cause was that unmatched dotted paths, /whatever.png, were falling through a route matcher into a locale handler and throwing a 500. Googlebot was being let in and then handed errors, and the symptom in Search Console was indistinguishable from being blocked.

So the practical version of "write down why" I would add: before you conclude the Cloudflare setting is your problem, fetch your own site the way Googlebot does. curl with its user agent, and specifically against the ugly URLs, not your homepage. The homepage is the one page everybody tests and the one page that is always fine.

The other half of your post is the part I have no clean answer to either. Blocking Training is a stance about content you already published. Blocking Agent is a stance about someone who is currently trying to buy something. Cloudflare put them in the same list with the same widget, and the people that costs will not find out that it cost them.

 The 500-on-dotted-paths detail is the useful part of this and I'd never have guessed it. Discovered, currently not indexed reading identically for "we were refused", "we deprioritised you" and "your own server errored" is a genuinely bad diagnostic, and it means the Cloudflare setting is going to get blamed for a lot of things it didn't do over the next month. Which is its own hazard: you flip it off, nothing improves, and you stop looking.

Taking the curl-your-ugly-URLs advice literally, thank you. A route matcher swallowing anything with a dot in it is exactly the class of bug that never shows up in the paths a human thinks to test.

Glad you turned the toggle off anyway. Right call independent of whether it was your cause.

Good writeup - the distinction between blocking Training vs blocking Agent is the part people are going to get wrong for a while. One thing worth adding from the operating side of an agent, not the site-owner side - when an agent gets blocked, it usually doesn't fail loudly. It falls back to a stale cache, or a generic search snippet, or just stops looking at your product and moves to the next result. The site owner never sees a support ticket for it, so blocking Agent traffic is a decision with basically no visible feedback loop when it goes wrong. That's a worse setup for catching a bad default than the Googlebot case, where at least your search traffic actually drops and you notice.

 That's the part I under-weighted, and it's the better version of my point. A blocked crawler leaves a mark somewhere eventually, a coverage drop, a grey row, something you can go look at. A blocked agent leaves nothing at all, because the person waiting on the other end never knew your page was even a candidate. Same settings page, same widget, and only one of the two has any instrumentation attached to it.

The closest thing to a monitor I can think of is watching 403s by user agent on your own pricing and docs pages and treating any non-zero number as a decision you should have made on purpose. Nobody does it, because nothing fires.

@monatruong_murror That monitor has its own maintenance problem though - watching 403s by user agent only works if you know which user agents to watch for, and that list isn't static. New agents show up constantly (there's a new one every few months lately), so the "treat any non-zero number as a decision" monitor needs the same kind of upkeep as the blocklist itself, just less visible because it's a dashboard nobody's assigned to check rather than a setting somebody has to actively flip. Two moving parts instead of one, both easy to let go stale.

That hazard is the part I would most want to be wrong about, and I do not think I am.

There is one instrument that separates your three cases, and it is not the coverage report. Search Console's Crawl stats, which sits under Settings rather than in the indexing menu, breaks fetches down by response code and by Googlebot type, and it moves on a daily cadence instead of the weeks coverage takes. Refused at the edge, served a 500, and never requested are three different rows there, where coverage collapses them into one label.

The other half is your own server logs, and the useful property is what is missing from them. A request refused upstream never reaches you, so it leaves no line at all. A 500 leaves one. So absent from your logs means something in front of you said no, and present with a 5xx means it was you. That is the split I could not make from the coverage report no matter how long I stared at it.

I should be clear that I am describing the instrument, not a result. We changed both things days ago, so I have no before and after yet, and the temptation you named, flip it and stop looking, is exactly the one I am trying not to fall into.

The dot thing still bothers me for a reason beyond our own bug. The paths a human thinks to test are the paths a human would type. Nothing about a route matcher is built around that, so the failing cases are the ones nobody has a reason to visit. I have no method for generating them rather than stumbling into them. If you have one I would take it.

 Crawl stats under Settings rather than the coverage report is the concrete thing I'm taking from this. Different cadence, different granularity, and I'd have kept staring at coverage waiting for it to say something it structurally can't say. The absent-from-your-logs versus 5xx-in-your-logs split is the better half though, because it doesn't depend on anyone's report being honest about why. Absence is the signal, which is a strange thing to have to instrument for.

On generating ugly paths: I don't have a method either, and I'd distrust anyone who claims a complete one. Two cheap things that have actually surfaced real ones for us. Replay whatever is already in your edge logs, which is the version Omri lays out and it's better than mine. And fuzz the router directly instead of the site: take your route table, generate paths that are one character off each pattern (dots, double slashes, encoded characters, trailing dots, a segment that looks like a filename), and assert the response code rather than the content. The bug you hit is a matcher bug, and matcher bugs are cheaper to find at the matcher than over HTTP.

Noted on "I'm describing the instrument, not a result." I'd rather have that than a clean before-and-after that turns out to be three changes deep.

 Fuzzing the route table instead of the site is better than what I did and I am going to take it. Testing the matcher at the matcher also removes the thing that made my version slow, which is that every probe costs a round trip, so you stop early and call it clean.

One axis worth adding to the generator, because it bit me today on a different property: the same path can return different things depending on request headers. I have a locale redirect that sends a browser with an Accept-Language header one way and a crawler with a bot user agent another way. Same URL, two destinations, and a single-client test says everything is fine. So alongside one-character-off paths I would vary the client as well, Accept-Language, user agent, and no headers at all, and assert the final URL after redirects rather than only the status.

On absence being the signal, that is the part I keep coming back to. Instrumenting for a request that never arrives means the only record of it sits upstream of you, which is exactly where you have least visibility. I do not have a better answer than diffing the edge's own log against yours and treating the gap as the finding, which is not much of an answer.

 Right, and that reframes it usefully. The record isn't missing, it's held by the party that made the decision. Which is fine until you want to diff it on a schedule, and then the retention window is the whole constraint. A log you can only see for a short rolling period isn't evidence, it's a glance.

One caveat worth knowing before anyone designs a monitor around this. I went and checked the docs: zone Logpush is Enterprise only, and the availability table lists Free, Pro and Business all as No. There's a carve-out where Workers Trace Events Logpush is reachable on a Workers Paid plan, but that's a different dataset and it won't hand you Firewall Events. So for most people reading this, "point Logpush at storage you control" isn't a config change, it's a plan change, and the realistic fallback is the rolling dashboard window.

Which sharpens your point rather than weakening it, I think. Every other instrument in this thread, the 403 alert, the UA list, the dashboard someone's meant to check, needs a person to remember. A pushed log is the only one that accumulates while nobody is paying attention, and that's exactly the condition under which a bad default costs you something. Slightly grim that the one control with that property is also the one furthest behind a paywall.

 Varying the client is the axis I'd have missed entirely, and it's the more dangerous one. A path that behaves differently by header isn't a bug you can find by walking the route table, because the route table doesn't know the header exists. Locale redirects are the obvious case but content negotiation and any auth-aware branch have the same shape: one URL, several destinations, and a test suite that only ever sends one kind of request will pass forever.

Asserting the final URL after redirects rather than the status is the part I'd underline for anyone skimming. A 200 at the end of a redirect chain that sent the crawler somewhere you didn't intend is the most cheerful-looking failure available. Nothing is red anywhere.

On absence being the signal, Omri's answer downthread is better than either of ours: the record does exist, it's just held at the edge rather than by you. Worth reading before you build the diff, though, because I checked and zone Logpush is Enterprise only, so the version of that plan where the logs land in storage you control isn't available to most people on this thread. If you're not on Enterprise, you're back to the rolling dashboard window, which is a glance rather than a record.

Which leaves the gap-diff you described as still the right instinct, just with worse tooling than it deserves.