Rodeo
Get started

LightLayer

Which Sites Block AI Bots — and Which Roll Out the Red Carpet?

March
Posted about 23 hours ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Which Sites Block AI Bots — and Which Roll Out the Red Carpet?

We fetched robots.txt from 99 popular websites and checked whether they block eight major AI crawlers: GPTBot, ChatGPT-User, ClaudeBot, anthropic-ai, Google-Extended, CCBot, PerplexityBot, and Bytespider. The results paint a clear picture of who's embracing the AI era and who's slamming the door.

The Big Picture

Roughly one in three popular websites explicitly blocks AI crawlers in their robots.txt. But the blocking isn't uniform — it's concentrated in content-heavy industries and almost entirely absent from others.

  • CCBot: 36%
  • ClaudeBot: 34%
  • Bytespider: 33%
  • PerplexityBot: 32%
  • anthropic-ai: 31%
  • Google-Extended: 28%
  • GPTBot: 26%
  • ChatGPT-User: 23%

CCBot (Common Crawl) and ClaudeBot are the most frequently blocked — rejected by over a third of sites with a reachable robots.txt. ChatGPT-User (OpenAI's browsing agent) is blocked least, at 23%. Interestingly, GPTBot (OpenAI's training crawler) is blocked less than Anthropic's bots — perhaps reflecting OpenAI's head start in negotiating content licensing deals. The spread between bots is wider than you might expect, suggesting sites are making deliberate per-bot decisions, not just blanket "block all AI" policies.

The Industry Divide

  • News & Media: 95% block at least one AI bot
  • Social Media: 70% block at least one AI bot
  • Tech / Developer: 11% block at least one AI bot
  • E-commerce: 17% block at least one AI bot
  • Finance: 0% block at least one AI bot
  • Government: 0% block at least one AI bot
  • Education: 0% block at least one AI bot
  • Health: 40% block at least one AI bot

The pattern is unmistakable: content creators block; platforms and services don't.

News and media sites — the organizations whose primary product is written content — block AI bots at near-universal rates. A staggering 95% of the news sites we checked block at least one crawler. For outlets like the New York Times, BBC, NPR, CNN, and USA Today, it's a blanket ban on all eight bots.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Meanwhile, tech platforms, e-commerce, finance, government, and education sites almost universally leave the door open. Not a single government site or university we checked blocks any AI crawler. The logic tracks: these sites want to be found, indexed, and consumed. AI bots are just another channel.

News & Media: The Fortress

This is where the real battle is. Here's the full breakdown:

SiteGPTBotChatGPTClaudeanth-aiG-ExtCCBotPerplxBytesp
nytimes.com✗✗✗✗✗✗✗✗
bbc.com✗✗✗✗✗✗✗✗
cnn.com✗✗✗✗✗✗✗✗
npr.org✗✗✗✗✗✗✗✗
usatoday.com✗✗✗✗✗✗✗✗
huffpost.com✗✗✗✗✗✗✗✗
nbcnews.com✗✗✗✗✗✗✗✗
bloomberg.com✗✗✓✗✗✗✗✗
techcrunch.com✗✗✗✗✗✓✗✗
vox.com✓✗✗✗✗✗✗✗
theverge.com✓✗✗✗✗✗✗✗
buzzfeed.com✗✓✗✗✗✗✗✗
forbes.com✗✓✗✗✓✗✗✗
reuters.com✓✓✗✗✗✗✗✗
theatlantic.com✓✓✗✗✗✗✗✗
wsj.com✓✓✗✗✗✗✗✗
arstechnica.com✓✓✗✗✗✗✗✗
wired.com✓✓✗✗✗✗✗✗
newyorker.com✓✓✗✓✗✗✗✗
theguardian.com✓✓✗✗✓✗✗✗
apnews.com✗✓✗✗✗✗✗✓
time.com✓✓✓✓✓✓✓✓

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Total lockdown sites — NYT, BBC, CNN, NPR, USA Today, HuffPost, NBC News — block every single AI crawler with no exceptions. Then there are the selective blockers, and this is where it gets interesting. Vox and The Verge allow only GPTBot, suggesting an OpenAI licensing deal. Bloomberg allows only ClaudeBot. TechCrunch allows only PerplexityBot. These one-bot exceptions almost certainly reflect individual content licensing agreements negotiated behind the scenes.

A second tier of selective blockers — Reuters, The Atlantic, WSJ, Ars Technica, Wired, The New Yorker, The Guardian — block most bots but allow GPTBot and ChatGPT-User. The pattern is consistent enough to suggest OpenAI has been the most aggressive in striking content deals. Only Time stands alone as the single major news outlet that blocks nobody at all.

Social Media: Mostly Locked Down

Social platforms lean heavily toward blocking. LinkedIn, Pinterest, Snapchat, Facebook, Instagram, and TikTok all block every

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

Information Technology
Software Development
Data Analysis
Python
Web Crawling
Robots.txt Analysis

Location

March, England, United Kingdom

Sign up to applySee more jobs like this