Mark RadarMARK RADAR
About
EN
Sign in

Cloudflare Launches New /crawl API for AI and RAG Data Collection

1 reports · First detected 2026-03-12 · Last active 2026-03-12

Cloudflare has long helped websites block malicious crawlers and is now offering a compliant content-extraction tool. The /crawl API, introduced in July 2025, streamlines the creation of AI training datasets and retrieval-augmented generation knowledge bases, reducing the cost of handling site navigation, content cleaning and format conversion in-house.

The new /crawl API is currently in open beta. Developers can crawl an entire website with a single API call and convert its pages into Markdown, or use Workers AI to produce JSON in a specified structure. The service also supports incremental updates, allowing RAG pipelines to retrieve only new or modified content and improving data-refresh efficiency.

All Coverage

1 original reports

The Backstory

The history behind this event
After this
Cloudflare Lets Websites Block AI Training While Keeping Search Visibilityfirst seen 2026-09-16 · 2 reports · similarity 0.81

Website operators have struggled to prevent their content from being used to train artificial intelligence models without also sacrificing traffic from traditional search. The conflict is most acute with mixed-use crawlers such as Googlebot, Applebot and Bingbot, which can serve both search indexing and AI training. The issue matters for publishers and other sites funded by advertising, subscriptions or referrals because AI-generated answers can replace direct visits, weakening the business models that finance online content.

Cloudflare launched its Disallow AI Training setting on September 15, 2026, separating controls for Search, Training and Agent traffic. Apple, Google and Microsoft have met or committed to its “Accountable” requirements, allowing their mixed-use crawlers to index participating sites while honoring no-training preferences; other training crawlers can be blocked. New ad-supported domains are offered presets that allow search, disallow training and block AI agents on pages carrying ads. Cloudflare said 17% of sites restrict training, while fewer than 1% block search; mixed-use crawlers account for 36.6% of verified crawler traffic.

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)