Skip to content

article · IS MY WEBSITE BLOCKING AI CRAWLERS

Check It Yourself InTen Minutes.

A good share of the sites I audit in The Valley block at least one of the crawlers that feed AI answers, and nobody in the building knows it happened. This is the check, in the order I run it, with the file paths so you can do it without me.

trusted by: UFC, Caesars Entertainment, Churchill Downs, Neighborly, Supercuts, City of Las Vegas, SDMI, Switch, Silverton, Panda Express, TAO, Ellie Mental Health, Dollar Loan Center, Dot Vegas Domains, HFC, Lithion Battery

UFC
Caesars Entertainment
Churchill Downs
Neighborly
Supercuts
City of Las Vegas
SDMI
Switch
Silverton
Panda Express
TAO
Ellie Mental Health
Dollar Loan Center
Dot Vegas Domains
HFC
Lithion Battery

Trusted by UFC, Caesars Entertainment, City of Las Vegas, Supercuts

Justin's expertise, responsiveness, and genuine investment in our success have been evident throughout this process. We are truly grateful.
Dr. Linda Silvestri & Dr. Angela Silvestri-Elmore, Co-authors, Saunders Pyramid to Success · 14 years · 300+ Las Vegas businesses · 5-star rated
drop: checking a website robots file for blocked ai crawlers

01 · what are ai crawlers

What AI crawlers are.

Separate from Google's normal crawler, there is now a second set of automated readers, one per AI product. They are the reason your business shows up when somebody asks ChatGPT or Perplexity or Google's AI answer who to call. If one of them cannot fetch your pages, you are not in the answer.

These readers do two different jobs, and the distinction is the whole decision later on. Some fetch pages to train a model. Some fetch pages to answer the question a person just typed, and cite the source. Blocking the first costs you nothing today. Blocking the second removes you from the answer.

02 · how to check if chatgpt can read my website

Where to check for a block.

Five places, in this order. The first two catch most of it and take about a minute each.

01

Your robots.txt file

Type your domain followed by /robots.txt into a browser. Read for any Disallow line that sits under a User-agent line naming GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Applebot-Extended or Google-Extended. A Disallow of a single forward slash means the whole site.

02

Your firewall or security plugin

Robots.txt is a request. A firewall is a wall. Look in Cloudflare's bot settings, or your security plugin's bot list, for anything switched on that reads as blocking AI bots or unverified bots. This is the most common cause I find, and it never shows up in robots.txt.

03

The page source, for meta tags

A noai, noimageai or a robots noindex tag in the head of a template quietly covers every page built from that template. Check the templates that matter, usually the home page and the service pages.

04

Whether the content needs JavaScript to appear

Turn JavaScript off in your browser and reload. If your headline and your body copy vanish, several of these readers see the same empty page. Nothing is blocked and you are still absent.

05

Your server logs

The one check that gives you a fact instead of an inference. Filter the last thirty days for those agent names. If a name never appears, either it never came or it never got through, and both are worth knowing.

03 · why would a website block ai

How the block usually happens.

Almost never on purpose. In the audits I publish, the block traces back to one of four things, none of them a decision anybody in the business made.

A security plugin shipped an update with a broader bot list than the version before it. A host turned on bot protection at the account level. A developer copied a robots.txt from a staging site where blocking everything was correct. Or somebody read an article about protecting content from AI and switched off the part that gets you recommended at the same time.

04 · should i block ai crawlers

Whether to block AI crawlers.

There is a defensible reason to opt out of model training. A firm with proprietary written work, or a photographer whose images are the product, has something real to protect, and the training opt-outs exist for exactly that.

There is almost never a reason to block the retrieval readers. Those fetch a page because a buyer asked a question this minute, and they cite what they fetch. Blocking them is the same decision as asking Google to remove you from search, made without anyone noticing they made it.

So the answer for most professional practices is split, not all or nothing. Opt out of training if the content is the asset. Stay open to retrieval always.

The pattern to hold on to: the tokens ending in Extended are training permissions, and the ones with Search or User in the name are how you get cited.
agent nameWho it belongs toWhat blocking it costs you
GPTBotOpenAI, crawling at scaleA training opt-out. Little to no visible cost today
OAI-SearchBotOpenAI, the ChatGPT search indexYou stop being a citable source inside ChatGPT
ChatGPT-UserOpenAI, fetching a page a user asked aboutThe one person actively looking at you gets nothing
ClaudeBot and Claude-UserAnthropicYou stop being citable in Claude
PerplexityBotPerplexityYou leave the answer engine most likely to send a click
Google-ExtendedGoogle, a training permission tokenA training opt-out only. It does not affect Google search
Applebot-ExtendedApple, a training permission tokenA training opt-out only. Apple search is unaffected

05 · how to get cited by ai

What makes a page quotable.

Access is the floor, not the finish. Once the readers can reach the page, the pages that get quoted have a shape in common, and you can check it by eye.

One question per heading, phrased the way a person would ask it out loud
The answer in the first two sentences under that heading, before any setup
Real specifics rather than adjectives. A number, a date, a place, a name
Prices and terms written on the page instead of behind a form
The same business name, address and phone on the site as on the Google Business Profile
Anything you claim, linked to a page where a stranger can see it for themselves

06 · ai crawler questions

AI crawler questions.

Does blocking GPTBot remove me from ChatGPT?
Not from the live answers. GPTBot is the bulk crawler associated with training. The agents that decide whether you appear as a cited source in ChatGPT are OAI-SearchBot and ChatGPT-User, and those are the ones worth leaving open.
Will allowing AI crawlers hurt my Google rankings?
No. Google-Extended is a permission token for Google's AI training and is explicitly separate from search. Allowing or disallowing it does not change how Googlebot crawls or ranks the site.
How do I know it is fixed?
Two ways. Your server log shows those agents fetching pages and getting a 200 response, and asking the assistants a question your business should own returns your pages. Check both. A robots.txt edit that a firewall still overrides looks fixed and is not.
Is llms.txt worth adding?
It costs nothing and no major search engine has committed to reading it. Add it if you like, after the access problem is closed. It is not a substitute for being fetchable.
Can you check this for me?
Yes, and the check is part of the free audit. I crawl the site, test what each of these readers can actually fetch, and send you the list of what is blocking what.

07 · ai search visibility las vegas

Where AI search visibility sits.

Being fetchable is one phase of the search work, not a project of its own. It lives beside the technical work, the pages themselves and the Google Business Profile, and I do all of it as one line.

01

AI search visibility

The phase that checks what the answer engines can read and cite.

02

Technical SEO

Crawlability, indexing and speed, checked line by line.

03

The written audit

Everything above, published at a URL you can open yourself.

I will tell you what is blocked.

I will crawl your site, test every one of these readers against it, and send you a ranked list of what is costing you visibility. Free. No call unless you want one.