---
title: "robots.txt"
path: https://www.eonshowroom.eon.com/en/aem-basics/robotstxt.html
lastModified: 2026-09-23T06:22:43.245Z
---

# robots.txt

# robots.txt

The **robots.txt**file controls which automated crawlers and search indexers may access your website. In addition to traditional search engines (Google, Bing), modern Generative Engine Optimization (GEO) requires intentional management of AI services such as OpenAI (ChatGPT / GPTBot), Anthropic (ClaudeBot), and Perplexity.

A proper configuration ensures public, commercial content remains crawlable and indexable by AI and search engines, while protecting internal system paths and staging environments.

### robots.txt Features

- Environment Control: Allows crawling on production while blocking non-production environments.
- AI & Search Readiness: Ensures AI agents (e.g. GPTBot, Perplexity) and search engines are not blocked.
- Path Protection: Excludes technical and internal directories (/system/, /libs/).
- Content-Signal: Defines content usage permissions for AI search context and model training.
- Sitemap Reference: Links crawlers directly to canonical XML sitemaps.
- Anchor ID support for internal linking

## Configuration

### robots.txt in a nutshell

Main rules for managing your website's crawler configuration:

- <strong>Production:</strong> Keep public content open (Disallow: empty or targeting only system paths).
- <strong>Non-Production:</strong>Block all crawler traffic globally (Disallow: /).
- <strong>AI-Agents:</strong> Ensure bots like GPTBot, ClaudeBot, and PerplexityBot are not blocked unless explicitly required by brand/legal policy.
- <strong>Content-Signal:</strong> Add standardized content-usage signals for search, context, and training.
- <strong>Location & Publication:</strong> Rich Text Editor for the response. Supports links and bullet points.

### In this how to

Quick overview of the most relevant how to topics

- Where to Maintain the robots.txt in AEM
- Production vs. Non-Production Configurations
- Content-Signal Directive (AI Usage Preferences)
- Verifying AI Crawler Access
- Step-by-Step Validation Guide

![](https://www.eonshowroom.eon.com/cdm-assets/is/content/eon/digit-1?ts=1789378602977)

#### Where to Maintain the robots.txt in AEM

Depending on your brand and solution setup, the robots.txt is maintained via one of the following mechanisms:

**1. Context-Aware Configuration (CAC) / Site Config (Standard):**

Configured at the root page level of your site (e.g. */content/<brand>/<country>/<lang>*). Site administrators maintain crawler directives and sitemap references directly via the Site Configuration dialog.

**2. Static Asset in DAM:**

In specific multi-brand setups, robots.txt is uploaded as a plain text file into the DAM and routed via Apache Dispatcher rewrite rules to the root domain ([https://www.yourdomain.com/robots.txt](https://www.yourdomain.com/robots.txt)).

**3. Dispatcher / CDN Rule:**

If your solution does not expose a root-level dialog, changes are managed via AEM Platform Support (AEMPS) deployment tickets.

**Note:** Updating the file or dialog on AEM Author has **no live effect** until it is replicated/published to the Dispatcher.

![](https://www.eonshowroom.eon.com/cdm-assets/is/content/eon/digit-2?ts=1789378603122)

#### Production vs. Non-Production Environments

Public production websites must allow search engines and AI assistants to discover brand content. Technical system folders are explicitly disallowed:

![](https://www.eonshowroom.eon.com/cdm-assets/is/image/eon/Screenshot%202026-09-17%20084359?qlt=50&fmt=webp-alpha&ts=1789628365114)

![](https://www.eonshowroom.eon.com/cdm-assets/is/content/eon/digit-3?ts=1789378603238)

#### Non-Production Websites (Stage, QA, Acceptance, Dev)

Staging environments (e.g. stage-www.eon.de) must **never** be indexed by search engines or scanned by AI scrapers:

![](https://www.eonshowroom.eon.com/cdm-assets/is/image/eon/Screenshot%202026-09-17%20084409%3A2-1?qlt=50&fmt=webp-alpha&ts=1789628859699)

(No sitemap or Content-Signal directives should be present on non-production instances).

##### Content-Signal Directive

Content-Signal is an emerging, standardized directive that communicates to AI providers how your public website content is permitted to be consumed.

![](https://www.eonshowroom.eon.com/cdm-assets/is/image/eon/Screenshot%202026-09-17%20084419?qlt=50&fmt=webp-alpha)

| **Signal** | **Permitted Values** | **Purpose** |
| --- | --- | --- |
| **search** | yes / no | Use content for direct search indexing and web search results. |
| **ai-input** | yes / no | Use content as real-time context / input for AI-generated answers (e.g. ChatGPT web browsing, Perplexity answers). |
| **ai-train** | yes / no | Use content to train or fine-tune underlying foundation AI models. |

**Validator Note:** Because Content-Signal is an emerging standard, some third-party robots.txt syntax validators may label it as "unknown directive". This is expected behavior; compliant crawlers simply ignore tags they do not yet support without breaking the remaining rules.

##### Verifying AI Crawler Access

Ensure your production robots.txt does not include unintentional crawler blocks. Common AI user-agents include:

- GPTBot (OpenAI / ChatGPT)
- ClaudeBot (Anthropic Claude)
- PerplexityBot (Perplexity AI)

Unfavorable / Blocking Example (Avoid on Production):

![](https://www.eonshowroom.eon.com/cdm-assets/is/image/eon/Screenshot%202026-09-17%20095326?qlt=50&fmt=webp-alpha)

Unless a specific legal or business directive dictates otherwise, remove bot-specific Disallow: / lines from production.

![](https://www.eonshowroom.eon.com/cdm-assets/is/content/eon/digit-4?ts=1789378603060)

#### Step-by-Step Validation Guide

- Access the live file: Open https://<your-domain>/robots.txt in an incognito browser window.
- Inspect root directives: Confirm that User-agent: * is not followed by Disallow: / on production.
- Verify AI bots: Check that GPTBot, ClaudeBot, and PerplexityBot are not listed under a Disallow: / block.
- Validate Sitemap: Ensure the Sitemap: URL points to a secure, valid https:// XML sitemap on the matching domain.
- Clear Dispatcher/CDN Cache: If edits were published but do not appear live, initiate an edge cache purge.

## Does, Don'ts

![](https://www.eonshowroom.eon.com/cdm-assets/is/image/eon/dcc25bcb-4735-4645-9822-b49fb3ae81b1?fmt=webp)

### What you should do

- Allow access to public production content: Ensure your public brand pages, products, and news articles remain fully crawlable.
- Align Content-Signal with corporate policy: Set search, ai-input, and ai-train strictly in accordance with approved corporate content guidelines.
- Use Disallow: / across all non-production instances: Prevent staging and QA environments from polluting search engine indices.
- Reference valid HTTPS sitemaps: Always include the canonical sitemap link at the bottom of the production robots.txt
- Perform a live browser check post-publication: Always verify the live file directly at /robots.txt after replicating changes.

![](https://www.eonshowroom.eon.com/cdm-assets/is/image/eon/9884279e-1390-4634-bda2-7b7d4dc1e149?fmt=webp)

### What you shouldn't do

- Do not use Disallow: / on production: A single root slash completely delists the entire website from search engines and AI engines.
- Do not block AI crawlers accidentally: Do not copy-paste legacy rules that block GPTBot or ClaudeBot without explicit stakeholder approval.
- Do not treat robots.txt as a security firewall: Disallowing a path does not prevent users or malicious bots from accessing sensitive URLs directly. Use AEM CUGs, ACLs, or SSO authentication instead.
- Do not publish staging configurations to production: Ensure deployment pipelines do not overwrite production robots rules with staging placeholders.

![](https://www.eonshowroom.eon.com/cdm-assets/is/image/eon/EON_Pattern_1-3%3A2-1?fmt=webp-alpha)
