# Nematix Portal - robots.txt # https://nematix.com/robots.txt # # --------------------------------------------------------------------------- # The policy, in one line: no training, yes to search and to answer engines # that cite us. # # Those are different things, and the previous version of this file did not # separate them — it blocked OAI-SearchBot (which builds ChatGPT's *search* # index) in the same breath as GPTBot (which builds OpenAI's *training* set). # The effect was to opt out of being cited by an AI search product while # leaving a dozen actual training crawlers unlisted and therefore allowed. # That is backwards for a site that publishes /.well-known/agent-skills/ and # an RFC 9727 API catalog specifically so that agents can read it. # # So: # - Answer and search engines that link back are allowed. They send traffic. # OAI-SearchBot, PerplexityBot, Claude-SearchBot, Claude-User, # ChatGPT-User, DuckAssistBot, YouBot, Amazonbot and plain Applebot all # fall through to the `*` group below and are welcome. # - Crawlers whose documented purpose is assembling a training corpus are # blocked by name. Being absent from that list is what "allowed" means, so # the list has to be reasonably complete or the policy is not a policy. # "Reasonably complete" is checkable rather than a matter of taste: Ahrefs # Site Audit reports allowed vs blocked training bots per page, and the # target is an empty "allowed" column. Matching there is case-insensitive # per RFC 9309 — `meta-externalagent` below is reported as blocking # `Meta-ExternalAgent` — so spelling a name in its documented casing is # safe. # # Content-Signal (below) states the same thing once, for every crawler, # including ones nobody has named yet. The per-agent Disallow blocks are # enforcement for the known training crawlers that do not read it. # /.well-known/agent-skills/index.json advertises this header as the canonical # statement of our content-usage preferences — keep the two in agreement. # Spec: https://developers.cloudflare.com/ai-crawl-control/features/managing-ai-crawlers/ # --------------------------------------------------------------------------- # --- AI training crawlers: blocked ----------------------------------------- # Documented purpose is building or selling a model-training corpus. None of # these surface a citation or a link back, so there is nothing on the other # side of the trade. # OpenAI — model training. (Its search and user-initiated agents are allowed.) User-agent: GPTBot Disallow: / # Google — Gemini training and app grounding. Does NOT affect Google Search or # AI Overviews, which crawl as ordinary Googlebot and remain fully allowed. User-agent: Google-Extended Disallow: / # Apple — Apple Intelligence training. Plain Applebot, which powers Siri and # Spotlight search, is a different agent and is allowed. User-agent: Applebot-Extended Disallow: / # Anthropic — general/training crawler. Claude-SearchBot and Claude-User are # separate agents that cite sources, and are allowed. User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / # Common Crawl — the corpus most public models are trained on. User-agent: CCBot Disallow: / # Meta — model training. Meta-ExternalFetcher (user-initiated link previews) # is a different agent and is allowed. User-agent: meta-externalagent Disallow: / User-agent: FacebookBot Disallow: / # ByteDance User-agent: Bytespider Disallow: / User-agent: TikTokSpider Disallow: / # DeepSeek — model training. # # Added 2026-09-08. The 2026-09-07 Ahrefs crawl reported this and xAI-Bot as # the only two training crawlers on its list still reaching us, across all 97 # indexable pages: six blocked, two allowed. A policy that is six-eighths # applied is the failure mode the header of this file describes — the list # going stale as new crawlers appear — so it is worth naming that the gap was # measured, not guessed. User-agent: DeepSeekBot Disallow: / # xAI — Grok's training crawler. # # xAI publishes no crawler documentation, so this token is the one Ahrefs # tracks rather than one xAI has confirmed. That is a deliberate limit: it # blocks a crawler identifying as exactly `xAI-Bot` and nothing else. Grok's # retrieval agents are reported to use different product tokens # (`xAI-Grok/1.0`, `Grok-DeepSearch/1.0`), which do not match this group and # keep falling through to `*` below — so this does not repeat the # OAI-SearchBot mistake of opting out of an answer engine that cites us. # If xAI ever documents a citing agent by name, it belongs in the allow list, # not here. User-agent: xAI-Bot Disallow: / # Cohere User-agent: cohere-ai Disallow: / User-agent: cohere-training-data-crawler Disallow: / # Data brokers that crawl to resell to model trainers. User-agent: Diffbot Disallow: / User-agent: omgili Disallow: / User-agent: omgilibot Disallow: / User-agent: Webzio-Extended Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Timpibot Disallow: / # Allen Institute for AI — the Dolma training corpus. User-agent: AI2Bot Disallow: / User-agent: AI2Bot-Dolma Disallow: / # Huawei User-agent: PanguBot Disallow: / User-agent: Kangaroo Bot Disallow: / User-agent: iaskspider/2.0 Disallow: / # --- Everyone else --------------------------------------------------------- # Ordinary search crawlers, and the AI answer engines that cite and link. User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / # Application surfaces with nothing to index. Disallow: /admin/ Disallow: /api/ Disallow: /user/ # The sales deck is an internal tool, not a page. It is already served # noindex (src/layouts/deck-layout.astro) and is linked from nowhere public; # this keeps it out of crawl budget and out of analytics too — the GA and # Ahrefs tags are gated off the same path in # src/components/layout/analytics.astro. Disallow: /deck/ # Sitemap location Sitemap: https://nematix.com/sitemap-index.xml