
Featured · client work
The major AI companies publish the names of their crawlers and state that those crawlers follow robots.txt. The standard is voluntary and unenforceable, researchers have documented crawlers that ignore it, and rules written for a training crawler often do not cover the fetcher a chatbot uses when a user asks about a specific page.
Training crawlers collect text to train future models. Search crawlers build the index an assistant queries when it answers. User-triggered fetchers retrieve one page because somebody pasted a link or asked about a specific site. Each has its own user-agent name, and blocking the first does nothing to the other two.
Blocking a search crawler removes a site from the pool an assistant draws on, which is the opposite of what most businesses want. Sites that blanket-blocked everything with AI in the name during 2024 removed themselves from AI answers while leaving their training exposure roughly unchanged.
It is a request, honoured by convention. It offers no protection against a crawler that renames itself or routes around the rules, and content already absorbed into a trained model cannot be withdrawn by adding a line to the file. Anything that genuinely must not be public belongs behind authentication, not behind a disallow rule.
Free written audit. No call required, no commitment, no upsell at the end.
Reply within two business days