clic para entrar
← back to blog
Software

How to let AI read your website: robots.txt and llms.txt

How to let AI read your website: robots.txt and llms.txt

To let AI read your website, you check your robots.txt file and make sure you’re not blocking AI crawlers like GPTBot, ClaudeBot, PerplexityBot, and Google-Extended; llms.txt is a newer, optional proposal that, today, the major search engines say they don’t use. In other words: the main door is robots.txt, and it’s up to you to leave it open.

Many business owners don’t know their site is already telling AI whether it can come in or not. That instruction lives in a tiny file called robots.txt, and sometimes it’s blocking AI without anyone consciously deciding so. It’s worth understanding, because whether ChatGPT, Perplexity, or Gemini can read and cite you depends on it.

What is robots.txt and what is it for?

It’s a text file that lives at the root of your site (at yourdomain.com/robots.txt) and tells automatic crawlers—the “bots” that roam the web—what they can enter and what they can’t. It has existed for decades and is the standard mechanism for granting or removing crawl permission.

Each bot identifies itself with a name, and in robots.txt you can allow or block each one. The most relevant AI crawlers today are:

  • GPTBot, from OpenAI (the company behind ChatGPT).
  • ClaudeBot, from Anthropic (the company behind Claude).
  • PerplexityBot, from the AI search engine Perplexity.
  • Google-Extended, the signal you use to control whether Google uses your content for its AI products (separate from normal search crawling).

If your robots.txt blocks any of them, that AI won’t be able to read your content and, therefore, will hardly be able to cite or recommend you.

Should I let AI in or block it?

It depends on your goal, and that’s precisely the dilemma. Blocking AI crawlers protects your content from being used without control, but at the same time it makes you invisible to those platforms. It’s a business decision, not just a technical one:

  • If you want to be cited and recommended by AI, let it in. For most small businesses seeking more visibility and customers, shutting the door on GPTBot or PerplexityBot means closing off a growing channel where people already ask questions and make decisions.
  • If your concern is protecting proprietary content—courses, research, exclusive material—it may make sense to block certain crawlers in those sections. But keep in mind that you’re also giving up appearing in those AIs’ answers.

For most businesses that want to grow, the balanced recommendation is to let AI crawlers in and focus on having well-made content, reserving blocks only for truly sensitive sections.

An important nuance: not all AI bots do the same thing. Some crawl your site to train models, others to answer questions in real time citing sources (like Perplexity), and others for both. If your goal is to appear and be cited when someone asks, you’re mainly interested in letting through the ones that aim to answer with sources. Telling them apart is a technical matter, but the underlying idea is simple: blocking blindly can pull you out of exactly the places where you’d want to appear.

A file with allow and deny toggle switches, and small crawler bots queued waiting to pass
You decide which AI crawlers can enter your site and which sections.

And llms.txt? Do I need one?

llms.txt is a newer, optional proposal to guide AI models, but today you shouldn’t expect results from it. The idea is interesting: a file where you tell AIs what your most important content is and how to understand your site, a kind of “map” designed for language models.

The problem is that it’s an emerging and debated standard. Google, through John Mueller and Gary Illyes (July 2025), noted that the big bots don’t use it today. That is: no one guarantees that adding it will do anything for now. That doesn’t make it useless—it may consolidate in the future—but it does help you calibrate expectations.

Our honest stance for a small business:

  • It’s not mandatory. Nothing bad happens if you don’t have an llms.txt. Your priority should be robots.txt and the quality of your content.
  • Implement it if you want, as a bet on the future. It’s low-cost and does no harm. If your developer adds it, go ahead; but don’t add it expecting an immediate visibility jump.
  • Don’t confuse it with access permission. llms.txt guides; robots.txt is what actually allows or blocks crawling. That’s the one you really should review.

robots.txt is the door; llms.txt is, for now, barely a sign that many still don’t read.

What should I actually do?

Start with what has a proven effect: review your robots.txt. These are the sensible steps for a small business that wants to be cited:

  • Open yourdomain.com/robots.txt in your browser and check whether any line blocks GPTBot, ClaudeBot, PerplexityBot, or Google-Extended. If you don’t know how to read it, ask your developer to check it.
  • If you’re unintentionally blocking AI and want to be cited, remove those blocks. Sometimes they come set by default in certain templates or plugins.
  • Reserve blocks for what truly must be protected, not for the whole site.
  • Leave llms.txt as optional, knowing its impact today is uncertain.

In summary

Controlling which AI reads your site is simpler than it seems: almost everything is decided in robots.txt, that small file that grants or removes permission to crawlers. For a small business that wants to be cited and recommended by ChatGPT, Perplexity, or Gemini, the normal thing is to let them in and focus on having good content. llms.txt is a promising but still uncertain proposal: implement it if you want, without expecting miracles.

If you’re not sure what your site is allowing or blocking today, at Normandia Web we can review your robots.txt, leave your site accessible to the right AI crawlers, and make sure nothing is shutting the door on you by accident.

Ready to put it to work in your company?

Tell us what’s costing you time, money or control. We’ll help you figure out where to start.

Start your consultation →