AI Spotlight · Infrastructure & Policy
AI Agent Crawlers Now Need Permission. Here's How to Get It
From September 15, Cloudflare will block AI agent crawlers by default on any page that runs ads. Here is what changes, why Google is caught in the middle, and what agent builders need to do before the deadline.
AI agent crawlers, the bots that fetch pages in real time on behalf of a person waiting for an answer, will be blocked by default on a large slice of the web from September 15 onward, according to AI News. Cloudflare announced the change on July 1, and most coverage since has fixated on Google. The more useful part of the announcement is what it asks of anyone building agents, and what it offers them in return.
Cloudflare has replaced its old single block-AI-bots switch with three categories, as detailed in its official press release. Search covers bots that index a page to answer questions about it later. Agent covers automated systems acting in real time for a user, including ChatGPT's fetch bot and browser-driving agents. Training covers crawlers that pull content into a model's weights. These controls went live for every customer, including the free tier, on July 1.
This issue breaks down what the new defaults actually mean, why Googlebot is a special case, and the concrete steps both agent builders and publishers need to take before September 15.
* * *
What's changing
Ads are now the dividing line between allowed and blocked
From September 15, the defaults change in a specific way: Training and Agent crawlers will be blocked on pages that display ads, while Search stays allowed. The new defaults apply to domains newly onboarding to Cloudflare, new sites set up by existing customers, and every existing free-tier customer, a detail most of the initial coverage skipped entirely [1]. Anyone who does not want the new defaults can opt out through their security settings before the date arrives.
Cloudflare's underlying logic is straightforward: an advertisement is evidence that a page was built for a human to land on. A search crawler that sends a reader back to the source is a referral. A bot that reads the page and hands the answer to someone else, without ever sending traffic back, is something else entirely, and that distinction is now enforced at the network level rather than left as a robots.txt suggestion a crawler can simply ignore.
|
The three new bot categories Search: indexes a page to answer questions later, stays allowed by default. Agent: acts in real time for a user, blocked by default on ad-supported pages. Training: pulls content into a model's weights, blocked by default on ad-supported pages. |
* * *
The Google problem
Why blocking Training also risks blocking Googlebot
There is a genuine complication baked into this system. Googlebot crawls for both search and training using a single, unified bot, so under the most restrictive setting, a site that blocks Training also blocks Googlebot, and with it, the site's search visibility. That is not a minor technical footnote; it is the single biggest reason most coverage of this announcement fixated on Google specifically [2].
Cloudflare CEO Matthew Prince has been candid about the intent behind this pressure. He said the company hopes the changes will encourage mixed-use crawlers to separate search from agent use and training, framing bot traffic surpassing human traffic online as the deeper trend forcing this reckoning sooner than expected [1]. That is, in effect, a polite way of saying the pressure on Google and other mixed-use crawlers to split apart is the entire point of the deadline, not a side effect of it.
For any company running a crawler that touches both search indexing and model training under one identity, September 15 is the forcing function to finally separate those functions technically, not just describe them differently in a policy document.
Granola Runs Revenue On Attio
"When I think of revenue, I think of Attio." - Shreman Shrestha, Head of Business at Granola
Here's what that adds up to:
Zero missed leads and 10x faster access to customer context
Lead triage 83% faster
Five hours saved per week with automated updates
If you build agents
The failure mode isn't a lawsuit. It's silence
Agentic deployments have largely been built on the assumption that the open web stays open. A research agent fetches a competitor's pricing page. A monitoring tool checks a supplier's announcements. A customer-service agent pulls a manufacturer's specification sheet. None of this has ever involved a licence, and until now, none of it needed one. Cloudflare sits in front of a large share of the world's web traffic, and its blocks operate at the network level, not as an easily ignored robots.txt suggestion.
That matters because ad-supported pages are exactly the pages agents want most. That is where news, reviews, pricing, and product coverage actually live. So the practical failure mode for an enterprise agent after September 15 is not a lawsuit or a takedown notice. It is silence, or an answer quietly built from whatever thinner set of sources the agent could still reach.
|
Before September 15, if you run agents Work out which of your Cloudflare accounts will read as Agent-class. Classification is behavioural, not opt-in, so a research agent that browses in real time gets caught regardless of what its operator calls it. Expect degraded coverage rather than a clean failure, since the block only lands on ad-supported pages. Negotiated access, not a rewritten user-agent string, is the actual way through. |
* * *
If you run a website
Free-tier publishers get switched over automatically
Publishers have a different homework list, and the first item is easy to miss. Check your Cloudflare tier first, since existing free-tier customers are moved to the new defaults automatically on September 15 without any action required on their part. Then decide whether blocking Training is actually worth what it costs, because as covered above, it takes Googlebot down with it, and your organic search visibility along with it.
The mechanism most worth watching here is the money. Pay Per Crawl is evolving into Pay Per Use, with Ceramic.ai paying publishers when their content appears in AI search results, and You.com paying when an agent reaches premium content [3]. Cloudflare says more than half of all AI crawler traffic is spent re-fetching pages that have not changed since the last crawl, which means there is real waste on both sides of this transaction worth pricing out properly.
| If you... | Check this before Sept 15 |
|---|---|
| Run agents that fetch live pages | Which accounts read as Agent-class, and negotiate access ahead of the deadline |
| Run an ad-supported website | Your Cloudflare tier, and whether blocking Training is worth losing Googlebot too |
| Run a mixed search/training crawler | Separate the crawler's identities before the deadline forces the issue |
* * *
The open question
The whole system runs on self-declared honesty
There is a real weakness sitting in the taxonomy itself. Search, Agent, and Training are behaviours that AI companies declare about their own bots, and a firm that would rather not have its training activity classified as Training has an obvious incentive to describe it differently. Cloudflare's announcement does not explain what actually stops that from happening at scale, beyond the threat of enforcement after the fact.
|
Access to the open web has been free and unlimited for thirty years, and the bill is now itemised. Agent builders who sort out their access before September have a workable problem. The ones who find out from a 403 will be rebuilding on the fly. |
* * *
|
Before you go If you run agents or a website, have you checked how Cloudflare's new categories classify your traffic yet? Hit reply with one sentence. The most interesting answers will shape a follow-up issue tracking how the September 15 deadline actually plays out. |
Until next time,
AI Spotlight
Practical AI, translated into real work, once a week.


