Cloudflare’s AI crawler changes take effect September 15, 2026, and every service business behind Cloudflare needs to make one decision before then: do not accidentally block the crawlers that feed AI recommendations. Cloudflare now sorts AI bots into three groups, Search, Agent, and Training, and on pages that serve ads it will block Agent and Training bots by default. Agent bots are the ones ChatGPT and Perplexity use to fetch your site in real time when they answer a question, so blocking them can quietly drop you out of AI answers. The fix is simple: open your Cloudflare dashboard before September 15, keep Search and Agent access on, and confirm your robots.txt is not waving AI tools away. Cited Co is a boutique AI visibility agency founded by Lauren Lerner in Scottsdale, Arizona, and this setting is one of the first things we check in a client audit. You can see how we approach the technical side on our services page.
Here is why this matters more than it sounds. AI tools do not recommend businesses they cannot read. When ChatGPT, Perplexity, Gemini, or Claude answers a question about who to hire, it often fetches live pages to ground that answer. Block the wrong bot and you become invisible at the exact moment a buyer is asking for a name.
What the Research Shows
The scale of AI crawling is the reason Cloudflare acted, and the numbers are worth reading before you touch a setting.
Cloudflare’s own attribution data, published in 2026, found crawl-to-referral ratios ranging from about 118 crawls per human referral at the low end to nearly 50,000 at the high end. In plain terms, some AI crawlers take content tens of thousands of times for every visitor they send back.
Cloudflare also reported in 2026 that more than half of the crawl traffic from bots it considers legitimate re-fetches pages that have not changed since the last visit, which is a lot of taking for very little giving.
According to Cloudflare’s 2026 announcement, the new system splits AI bots into three named categories, Search, Agent, and Training, and starting September 15, 2026 it blocks Agent and Training bots by default on ad-serving pages while leaving Search allowed.
Cloudflare’s 2026 Content Signals policy also extends robots.txt with a new use directive, and customers on managed robots.txt now have use=reference appended automatically, which tells AI tools they may index and link but not reproduce wholesale.
Cloudflare CEO Matthew Prince framed the move in 2026 by noting that the majority of traffic on the internet is now non-human, which is the backdrop for treating model training and search indexing as different transactions.
As TechCrunch reported in 2026, the sharp edge for local businesses is Google: Googlebot bundles search indexing and training collection into one crawler, so blocking Training can put your Google visibility at risk until Google separates the two.
The Three Kinds of AI Crawlers, and Why the Difference Matters
Not all bots are the same, and the whole change hinges on telling them apart. Cloudflare’s three categories each do a different job, and each one affects your visibility differently.
Search crawlers index your pages so they can be surfaced later. Agent crawlers act in the moment, fetching your page because a user just asked an AI assistant something and the assistant is pulling live information. Training crawlers collect content to train future models. For a service business chasing AI recommendations, the Agent category is the one to protect, because that is the live fetch behind a real answer.
Here is the quick reference:
| Crawler type | What it does | September 15 default on ad pages | Why it matters to you |
|---|---|---|---|
| Search | Indexes pages for later results | Allowed | Feeds traditional and AI search listings |
| Agent | Fetches your page live when a user asks an AI assistant | Blocked | Directly powers AI recommendations, protect this one |
| Training | Collects content to train models | Blocked | Builds long-term model knowledge of your brand |
My honest recommendation for most service businesses is to keep Search and Agent open and treat Training as a business decision. If you are not a large publisher living off ad revenue, blocking Training mostly costs you future model familiarity while gaining you very little.
What the Cloudflare AI Crawler Changes Actually Do on September 15
The default flips, but only in specific cases, and knowing whether it applies to you takes two minutes.
The new blocking default hits pages that serve ads, and it applies automatically to new domains, new customers, and existing Free tier accounts. If you are on a paid Cloudflare plan, the defaults do not silently change your settings, but Cloudflare still recommends you review your preferences before the date. Either way, there is an opt-out window in your Security settings before September 15.
So the action is not complicated. Log in, find the AI crawler controls under Security, and confirm that Search and Agent are allowed for your site. If you run ads and you are on the Free tier, this is where the default could work against your AI visibility, so check it deliberately rather than assuming.
The Googlebot Trap Service Businesses Need to Avoid
This is the part that trips people up, and it is the reason a careless block can hurt more than help.
Google uses one crawler for both search indexing and training data. Because Cloudflare applies the most restrictive matching rule to a multi-purpose crawler, telling Cloudflare to block Training can also choke the crawler that keeps you in Google. For a local service business, losing Google visibility to protect training data is a bad trade. Until Google splits these functions, the safe play is to avoid blanket Training blocks unless you have a specific reason and you understand the Google cost.
You can verify how AI tools currently see you before and after you touch anything. Our guide on how to check your AI visibility score walks through the exact checks, and if you want the mechanics of getting quoted in the first place, how to write content cited by AI tools covers that side.
A Real Example
We watch this closely because we have seen how fast AI visibility can move when the plumbing is right. Living with Lolo, a Scottsdale interior design and construction firm, climbed from near invisibility to being cited as a top design-build firm across ChatGPT, Perplexity, Gemini, and Claude over a matter of months. None of that happens if the crawlers cannot reach the pages. A single misconfigured block would have undone the content and schema work behind those gains. You can read the full story on our case studies page.
The lesson generalizes. Great AI-ready content only helps if the AI tools are allowed to fetch it. Access is the floor, and content is the building you put on top.
What to Do Before September 15
Start with access, then move to signals. First, confirm in Cloudflare that Search and Agent crawlers are allowed, especially if you run ads or sit on the Free tier. Second, open your robots.txt and make sure you are not blocking the AI user agents you actually want, since a use=reference signal invites citation while a hard block ends the conversation. Third, decide your Training stance on purpose, with the Google bundling risk in mind. If this reads like a lot to weigh, that is fair, and it is the sort of thing an AI visibility agency exists to handle. Our explainer on what generative engine optimization is puts this change in context, and getting cited across ChatGPT, Perplexity, and Gemini shows what comes after access is sorted.
The businesses that win the next year of AI search in Scottsdale and across Phoenix will be the ones that treated crawler access as a real setting, not an afterthought.
Frequently Asked Questions
If I block AI crawlers in Cloudflare, will ChatGPT stop recommending my business?
It can, yes. ChatGPT and other assistants often use Agent crawlers to fetch your live pages when they answer a question, so blocking Agent access can remove you from those answers. After Cloudflare’s September 15, 2026 change, Agent bots are blocked by default on ad-serving pages, so a business running ads on the Free tier could lose AI recommendations without realizing it. Keep Search and Agent allowed to stay in the running.
Does Cloudflare’s September 15 change affect my site if I do not run ads?
Mostly no. The new default blocking applies to pages that serve ads, so a service site with no advertising is not the primary target of the change. That said, new domains, new customers, and Free tier accounts should still confirm their AI crawler settings, because defaults can shift under you. Two minutes in the Security settings is cheap insurance for your AI visibility.
What is the difference between an AI search crawler and an AI training crawler?
A search crawler indexes your pages so they can be surfaced in results later, while a training crawler collects content to train future AI models. Cloudflare added a third type in 2026, the Agent crawler, which fetches your page live when a user asks an assistant a question. For AI recommendations, the Agent crawler is the one that matters most, because it powers the answer in real time.
Will blocking AI training crawlers hurt my Google ranking?
It can, because of how Google is built. Google uses a single crawler for both search indexing and training collection, and Cloudflare applies the most restrictive rule to that combined crawler. So a Training block can also reduce the access Google needs to keep you in search results. Until Google separates these jobs, blanket Training blocks are risky for local businesses that depend on Google.
How do I check whether my site is blocking the AI crawlers that matter?
Start in your Cloudflare dashboard under Security, where the AI crawler controls show which of the three categories are allowed or blocked. Then review your robots.txt for any rules aimed at AI user agents like GPTBot or PerplexityBot. Finally, test how the assistants actually describe you, which our free snapshot tool does across ChatGPT, Perplexity, Gemini, and Claude in one pass.
Should a local service business block AI crawlers at all?
Usually not the Search and Agent types. For a Scottsdale or Phoenix service business trying to show up in AI answers, blocking those two works against you, since they are how assistants find and quote you. Training is a judgment call, and the honest answer for most small service brands is to leave access open and focus energy on being worth citing instead.
Cited Co is a boutique AI visibility agency founded by Lauren Lerner in Scottsdale, Arizona. We specialize in generative engine optimization for service businesses: entity optimization, schema implementation, AI-cited content creation, and monthly visibility tracking across ChatGPT, Perplexity, Gemini, and Claude. We work with a small, intentional roster of clients in high-consideration service categories where reputation drives decisions and AI visibility is becoming a meaningful competitive advantage.
Get Your Free Snapshot to find out where your business stands today.