Robots.txt rules for AI bots: agent groups in SEO/GEO Files Generator

Paper robots queuing in front of a signpost – some pass through an open gate, some are stopped by a barrier – illustration of robots.txt rules for AI bots
Summarize with AI:

Robots.txt is a file in a domain’s root directory with instructions for the robots that visit a website. Today, those robots also include crawlers run by companies developing AI models. BeeClear is a marketing agency working on search engine and AI-answer visibility, and it also builds its own WordPress plugins. SEO/GEO Files Generator is a BeeClear plugin that generates llms.txt, llms-full.txt and robots.txt files from a single panel. Its strongest feature in this area is rule groups for selected bots, backed by a database of 62 described user agents. See how to set up separate rules for AI crawlers without manually editing files on the server.

Why is one robots.txt rule for every bot not enough?

One robots.txt rule for every bot treats a search engine robot, a crawler harvesting data to train models, and an audit tool all the same way. Yet each of them visits your site for a different purpose. A “User-agent: *” entry covers everyone at once and lets you make no distinction at all. You can then either let everyone in or block everyone. There’s no middle ground where a search engine gets full access while a training crawler gets none.

Writing separate blocks by hand can be a hassle, too. You need to know the exact user-agent names, watch out for typos, and edit the file on the server. Any mistake in a name means the instruction simply never reaches its target. SEO/GEO Files Generator moves that work into the WordPress panel and offers ready-made names.

How do the AI bots visiting your site differ from each other?

AI bots visiting your site differ above all in their job. Some gather content to train models, others index pages for search inside AI assistants. Others still only fetch a URL when a chat user explicitly asks for it. The descriptions in the plugin let you recognize these roles without digging through every vendor’s documentation. Here are some examples from the built-in database:

Bot roleExample user agentsWhat it does, per the plugin’s description
model training and datasetsGPTBot, CCBot, Ai2Bot-Dolma, cohere-aigathers content to train models or build public datasets
AI searchOAI-SearchBot, Claude-SearchBot, PerplexityBotindexes content for AI search and answers
user requestsChatGPT-User, Claude-User, Perplexity-Userfetches a page on behalf of a person using the assistant
content-usage control tokensGoogle-Extended, Applebot-Extendedcontrols content use in Gemini and Vertex AI models and in Apple’s generative AI features
search engines and search servicesGooglebot, bingbot, Applebotindexes pages for search engines and search services

This breakdown leads to a practical conclusion. For some providers, separate user agents handle model training and AI search. So you can block one and leave the door open for the other. That kind of split is hard to achieve with a single block for every robot, but easy with a few groups in the plugin.

What’s in the plugin’s database of 62 user agents?

The database of 62 user agents in SEO/GEO Files Generator covers the crawlers and fetchers of the major providers. You’ll find bots from OpenAI, Anthropic and Perplexity, plus the crawler families of Google, Microsoft, Amazon and Apple. There are also other AI crawlers, SEO tools such as Ahrefs, Semrush, Majestic and Screaming Frog, and Meta’s fetchers. Every entry on the “Robots agents” screen has a short description of its role.

You can freely edit the list and add your own names, one per line. The plugin flags built-in entries that are missing from your list. A restore button adds them back without touching your own entries. When a new version expands the database, the new names appear once. Names you’ve deliberately removed don’t come back the next time the panel loads.

How do you build a rule group for selected AI bots?

You build a rule group for selected AI bots on the robots.txt screen, in the groups section. Each group is a set of checked bots plus a field for the directives that apply to them. You can add as many groups as you need. Here’s the full procedure:

  1. click “Add group” to add a new group,
  2. check the bot names in the grid that should follow the same rules,
  3. enter the group’s rules, for example “Disallow: /” or “Allow: /public/”,
  4. save the settings and check the result in the file preview.

For each group, the plugin writes a separate block with a “User-agent” line for every checked bot. It places the entered rules underneath them. A group with no bots checked, or with no directives, never makes it into the file, so an empty draft won’t break anything. In practice, you might, for example, block training bots while leaving access open for AI search bots.

How do global rules work together with bot groups?

Global rules go into the “User-agent: *” block and apply to every robot that has no group of its own. By default, they block the /wp-admin/ directory and let admin-ajax.php through. Groups are added below as separate blocks for specific bots. That way you set one baseline policy and spell out exceptions only where they make sense.

Between the global settings and the groups you’ll also find an extra content field. Whatever lines you enter there, the plugin inserts into the file unchanged. It’s the place for unusual directives you don’t want tied to a specific group. You still edit everything in a single form.

At the end of the file, the plugin adds “Sitemap” lines. In the field you supply just the paths, and the file gets the full addresses with the domain the site was opened under. When WordPress’s built-in sitemaps are turned off, the plugin skips the line pointing to them. So robots.txt never points to an address that ends in a 404.

How do you check robots.txt before exposing it to bots?

You can check robots.txt before publishing it in the preview next to the settings form. It shows exactly the text the plugin will generate. One click also opens the live file at /robots.txt. So you see every group in full context before any bot ever reads it.

You also get to choose how it’s published. WordPress can serve the generated robots.txt dynamically, or save it as a physical file in the root directory. A status line shows whether WordPress has write permission for that file. When writing isn’t possible, you can download the finished content with a button and upload it manually.

How does robots.txt connect with the Content-Signal directive?

The Content-Signal directive goes into robots.txt once you set it up in the BeeClear WebMCP AI Visibility plugin. Content-Signal is an entry declaring how automated systems may use a site’s content. SEO/GEO Files Generator appends it at the end of the file along with comments. They explain three signals: search for the search index, ai-input for AI answers, and ai-train for model training. The comments also note that restrictions expressed this way reserve rights under Article 4 of EU Directive 2019/790.

The plugin also makes sure this entry never appears twice. You can view the value on the robots.txt screen, and change it in the AI Visibility settings. Rule groups say who can come in, while Content-Signal says how they may use the content. Together, the two give a fuller picture of your policy toward AI.

How do you start organizing AI bot access to your site?

Start organizing AI bot access by enabling the robots.txt module in the plugin’s global settings. Then look through the “Robots agents” screen and the bot descriptions in the built-in database. Decide which bots you want to give content to and which to block. Finally, set up your groups and check the result in the preview.

We cover all the modules in detail on the SEO/GEO Files Generator page. You’ll also find screenshots of the robots.txt editor and the agent database there. If you’d rather have help with the rollout, the BeeClear team can help set up the plugin and configure llms.txt and robots.txt for your site. It only takes a few groups to give every bot clear rules instead of one rule for all.

Summarize with AI:

Want clients and AI models to find your website?

We run website SEO for Google and for AI answers. We tidy structure, internal linking and content, and we measure results instead of promising them.