# Sitemap Sitemap: https://sonomos.ai/sitemap.xml # ----------------------------------------------------------------------------- # AI crawler policy: no training, yes citations # ----------------------------------------------------------------------------- # Crawlers that collect training data are blocked. Crawlers that put Sonomos # into search results and AI answers are allowed. These are different # user-agents, so the two goals do not trade off against each other: # # - Blocking GPTBot does not remove us from ChatGPT search. That is # OAI-SearchBot (allowed below), and ChatGPT-User for user-triggered fetches. # - Blocking ClaudeBot does not remove us from Claude's citations. Those are # Claude-User and Claude-SearchBot (allowed below). # - Google-Extended governs Gemini training and grounding only. It has no # effect on ordinary Google Search inclusion, which Googlebot controls and # which stays allowed below. # - Applebot-Extended is Apple's training opt-out. Applebot itself (Siri, # Spotlight) is not blocked and falls under the default group. # # For a company selling a privacy layer, feeding LLM training crawlers is a # posture we would rather not have to explain. Blocking them costs us close to # nothing in AI-answer visibility. # # ----------------------------------------------------------------------------- # This file is NOT the whole response as served # ----------------------------------------------------------------------------- # Cloudflare prepends a "Cloudflare Managed content" block at the edge # (dashboard -> AI Crawl Control). That block sets # Content-Signal: search=yes,ai-train=no,use=reference for all crawlers and # sends Disallow: / to Amazonbot, Applebot-Extended, Bytespider, CCBot, # ClaudeBot, CloudflareBrowserRenderingCrawler, Google-Extended, GPTBot and # meta-externalagent. Fetch https://sonomos.ai/robots.txt to see the served # result; this file is only the second half of it. # # The groups below are written to AGREE with that managed block rather than # contradict it. Where two groups name the same user-agent, parsers disagree # about which one wins: Google merges same-name groups and breaks an # Allow/Disallow tie in favour of Allow, while a first-match parser stops at # Cloudflare's Disallow and never reads ours. Agreement is the only way to get # one answer out of every crawler. This file previously said Allow: / for # GPTBot, ClaudeBot, Google-Extended and Applebot-Extended while the managed # block said Disallow: / -- four contradictions whose real-world outcome # depended on whose parser was reading. # # One ambiguity is left and cannot be closed from this file: Cloudflare's # managed block opens with User-agent: * / Allow: /. That merges with the # default group below, so for a crawler with no group of its own, the path # restrictions (/admin/, /auth, /login, /checkout, /join/ ...) win under a # longest-match parser but lose under a first-match parser that stops at # Cloudflare's Allow: / and never reads ours. Verified: every named crawler # resolves identically under both semantics; only the unnamed-crawler case # diverges. The exposure is negligible -- those routes are auth-gated and carry # a noindex meta tag, and /admin and /dashboard do not serve content to an # unauthenticated fetch -- but closing it properly means editing the managed # block in the Cloudflare dashboard, not this file. # # If the managed block is changed or disabled in the Cloudflare dashboard, # revisit the training-crawler section below so the two stay in step. # ----------------------------------------------------------------------------- # Default + traditional search engines - allowed # ----------------------------------------------------------------------------- # No "Allow: /" here on purpose. robots.txt permits everything that is not # disallowed, so an explicit Allow: / adds nothing -- except an Allow/Disallow # tie against the path rules below, which longest-match parsers (RFC 9309, # Google) and first-match parsers resolve differently. Stating only the # restrictions gives every parser the same answer. # ----------------------------------------------------------------------------- User-agent: * Disallow: /admin/ Disallow: /account/ Disallow: /dashboard/ Disallow: /auth Disallow: /login Disallow: /checkout Disallow: /checkout-success Disallow: /checkout-cancel Disallow: /join/ User-agent: Googlebot Disallow: /admin/ Disallow: /account/ Disallow: /dashboard/ Disallow: /auth Disallow: /login Disallow: /checkout Disallow: /checkout-success Disallow: /checkout-cancel Disallow: /join/ User-agent: Bingbot Disallow: /admin/ Disallow: /account/ Disallow: /dashboard/ Disallow: /auth Disallow: /login Disallow: /checkout Disallow: /checkout-success Disallow: /checkout-cancel Disallow: /join/ User-agent: Yandex Disallow: /admin/ Disallow: /account/ Disallow: /dashboard/ Disallow: /auth Disallow: /login Disallow: /checkout Disallow: /checkout-success Disallow: /checkout-cancel Disallow: /join/ User-agent: DuckDuckBot Disallow: /admin/ Disallow: /account/ Disallow: /dashboard/ Disallow: /auth Disallow: /login Disallow: /checkout Disallow: /checkout-success Disallow: /checkout-cancel Disallow: /join/ # ----------------------------------------------------------------------------- # AI search and answer engines - allowed (these produce citations, not training) # ----------------------------------------------------------------------------- # LLM-readable summary: https://sonomos.ai/llms.txt # LLM-readable full text: https://sonomos.ai/llms-full.txt # LLM-readable HTML: https://sonomos.ai/llm/ User-agent: OAI-SearchBot Disallow: /admin/ Disallow: /account/ Disallow: /dashboard/ Disallow: /auth Disallow: /login Disallow: /checkout Disallow: /checkout-success Disallow: /checkout-cancel Disallow: /join/ User-agent: ChatGPT-User Disallow: /admin/ Disallow: /account/ Disallow: /dashboard/ Disallow: /auth Disallow: /login Disallow: /checkout Disallow: /checkout-success Disallow: /checkout-cancel Disallow: /join/ User-agent: PerplexityBot Disallow: /admin/ Disallow: /account/ Disallow: /dashboard/ Disallow: /auth Disallow: /login Disallow: /checkout Disallow: /checkout-success Disallow: /checkout-cancel Disallow: /join/ User-agent: Claude-User Disallow: /admin/ Disallow: /account/ Disallow: /dashboard/ Disallow: /auth Disallow: /login Disallow: /checkout Disallow: /checkout-success Disallow: /checkout-cancel Disallow: /join/ User-agent: Claude-SearchBot Disallow: /admin/ Disallow: /account/ Disallow: /dashboard/ Disallow: /auth Disallow: /login Disallow: /checkout Disallow: /checkout-success Disallow: /checkout-cancel Disallow: /join/ # ----------------------------------------------------------------------------- # AI training crawlers - blocked (matches the Cloudflare managed block above) # ----------------------------------------------------------------------------- # Anthropic-AI and Claude-Web are legacy Anthropic agent names, kept here so an # older crawler honouring the historical name gets the same answer. User-agent: GPTBot Disallow: / User-agent: GPTBot-Site Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-Web Disallow: / User-agent: Anthropic-AI Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: /