User-agent: * Allow: / # Pace generic/unnamed crawlers so a single bot can't walk the entire dynamic # /ar/* locale in seconds and spike Vercel Fluid CPU. Named AI bots that drive # citations (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, etc.) have their # own blocks below WITHOUT a crawl-delay — they are revenue-relevant and left # unthrottled by design. Googlebot ignores Crawl-delay. Bingbot does NOT: Bing # honors this default Crawl-delay even though Bingbot has its own group below # (blogs.bing.com/webmaster May-2012), so Bingbot is paced at 1 request / 10s. Crawl-delay: 10 # Content Signals (contentsignals.org / AIPREF) — machine-readable AI-usage # preferences. We welcome search indexing, AI answer/citation, and training on # our publicly served (Layer 1) content. The "ai-train=yes" signal carries an # attribution requirement this header cannot express — see /llms.txt and # /.well-known/agent.json for the binding terms (cite by name + link, preserve # footnotes). Commercial-scale access should use the Agent API — see /for-agents. Content-Signal: search=yes, ai-input=yes, ai-train=yes # Block API endpoints from indexing # (Agent API at /api/agents/v1/* is intentionally gated — see /for-agents for access) Disallow: /api/ # Block internal/utility routes Disallow: /_next/ Disallow: /404 Disallow: /500 # Block authentication pages Disallow: /auth/ Disallow: /login Disallow: /signup # Job detail pages are aggregated listings — Google prefers the original source # for ranking, so we take them out of the crawl path to preserve crawl budget # for first-party content (posts, courses, guides). /jobs index stays allowed. # See docs/seo-2026-04/SEO_ROOT_CAUSE_AUDIT.md (RC1). Disallow: /jobs/ Disallow: /ar/jobs/ # --- AI crawlers: explicitly allowed for Layer 1 (billboard) content --- # See /llms.txt and /.well-known/agent.json for full terms. # We welcome citation with attribution. Commercial-scale agent platforms should # use the paid Agent API at /for-agents instead of crawling. User-agent: GPTBot Allow: /posts/ Allow: /courses/ Allow: /guides/ Allow: /tools/ Allow: /podcast/ Allow: /llms.txt Allow: /llms-full.txt Allow: /.well-known/ Disallow: /api/ User-agent: ClaudeBot Allow: /posts/ Allow: /courses/ Allow: /guides/ Allow: /tools/ Allow: /podcast/ Allow: /llms.txt Allow: /llms-full.txt Allow: /.well-known/ Disallow: /api/ User-agent: Claude-Web Allow: /posts/ Allow: /courses/ Allow: /guides/ Allow: /tools/ Allow: /podcast/ Allow: /llms.txt Allow: /llms-full.txt Allow: /.well-known/ Disallow: /api/ User-agent: PerplexityBot Allow: /posts/ Allow: /courses/ Allow: /guides/ Allow: /tools/ Allow: /podcast/ Allow: /llms.txt Allow: /llms-full.txt Allow: /.well-known/ Disallow: /api/ User-agent: Google-Extended Allow: /posts/ Allow: /courses/ Allow: /guides/ Allow: /tools/ Allow: /podcast/ Allow: /llms.txt Allow: /llms-full.txt Allow: /.well-known/ Disallow: /api/ # A Bingbot-specific group makes Bingbot IGNORE every directive in the # User-agent: * group above except Crawl-delay (Bing, Dubut 2019: "You MUST # copy-paste the directives you want Bingbot to follow under its own section"). # Keep these Disallows in sync with the * group. /_next/ is deliberately NOT # copied: Bing is the site's main search channel and must be able to fetch CSS/JS. User-agent: Bingbot Allow: / Disallow: /api/ Disallow: /404 Disallow: /500 Disallow: /auth/ Disallow: /login Disallow: /signup Disallow: /jobs/ Disallow: /ar/jobs/ User-agent: Applebot-Extended Allow: /posts/ Allow: /courses/ Allow: /guides/ Allow: /tools/ Allow: /podcast/ Allow: /llms.txt Allow: /llms-full.txt Allow: /.well-known/ Disallow: /api/ User-agent: cohere-ai Allow: /posts/ Allow: /courses/ Allow: /guides/ Allow: /tools/ Allow: /podcast/ Allow: /llms.txt Allow: /llms-full.txt Allow: /.well-known/ Disallow: /api/ User-agent: FacebookBot Allow: /posts/ Allow: /courses/ Allow: /guides/ Allow: /tools/ Allow: /podcast/ Allow: /llms.txt Allow: /llms-full.txt Allow: /.well-known/ Disallow: /api/ # OpenAI's live ChatGPT-search retrieval bot (distinct from training-only # GPTBot) — it fetches pages to ground/cite answers, same citation value as # PerplexityBot. Unthrottled by design like the other named citation bots. User-agent: OAI-SearchBot Allow: /posts/ Allow: /courses/ Allow: /guides/ Allow: /tools/ Allow: /podcast/ Allow: /llms.txt Allow: /llms-full.txt Allow: /.well-known/ Disallow: /api/ # Meta's GenAI crawler (Llama / Meta AI citations). FacebookBot above is the # legacy link-preview UA — meta-externalagent is what Meta AI actually uses. User-agent: meta-externalagent Allow: /posts/ Allow: /courses/ Allow: /guides/ Allow: /tools/ Allow: /podcast/ Allow: /llms.txt Allow: /llms-full.txt Allow: /.well-known/ Disallow: /api/ Sitemap: https://nerdleveltech.com/sitemap.xml