# https://www.spekboom.org robots.txt # =========================================== # Default rules for all crawlers # =========================================== User-agent: * # Disallow admin, API, and internal routes Disallow: /api/ Disallow: /admin/ Disallow: /manage/ Disallow: /claim/ Disallow: /_next/ Disallow: /static/ # Allow important public pages Allow: /about # MCP connector docs — keep discoverable so the Spekboom AI connector # (Claude, ChatGPT, Perplexity) surfaces in search and AI answers. Allow: /mcp # Developer & agent resource hub (MCP server, OAuth scopes, llms.txt index) Allow: /developers Allow: /our-story Allow: /contact Allow: /places/ Allow: /search Allow: /trust Allow: /causes/ Allow: /explore/ Allow: /guides/ Allow: /stay/ Allow: /ask-dassie Allow: /plan-a-trip Allow: /regenerative-travel Allow: /gobirding Allow: /faq Allow: /help # =========================================== # Sitemaps # =========================================== Sitemap: https://www.spekboom.org/sitemap.xml Sitemap: https://www.spekboom.org/image-sitemap.xml # =========================================== # LLM Context Files # =========================================== # llms.txt: https://www.spekboom.org/llms.txt # llms-full.txt: https://www.spekboom.org/llms-full.txt # =========================================== # Search Engine Crawlers # =========================================== # Google # NOTE: a specific user-agent group REPLACES the * group for that bot, so the # outbound-handoff redirects (/api/go/) must be disallowed here explicitly — # they exist to keep link equity on-site, not to be crawled. User-agent: Googlebot Allow: / Disallow: /api/go/ # Bing / Microsoft Copilot User-agent: Bingbot Allow: / Disallow: /api/go/ # Apple (Siri, Spotlight) User-agent: Applebot Allow: / Disallow: /api/go/ # =========================================== # AI/LLM Crawlers — Allowed # We welcome AI crawlers to help users discover Spekboom # =========================================== # OpenAI GPTBot (training + search) User-agent: GPTBot Allow: / Disallow: /claim/ # OpenAI Search (ChatGPT search results) User-agent: OAI-SearchBot Allow: / Disallow: /claim/ # ChatGPT live browsing User-agent: ChatGPT-User Allow: / Disallow: /claim/ # Anthropic ClaudeBot (training) User-agent: ClaudeBot Allow: / Disallow: /claim/ User-agent: anthropic-ai Allow: / Disallow: /claim/ # Anthropic Claude search User-agent: Claude-SearchBot Allow: / Disallow: /claim/ # Anthropic Claude live browsing User-agent: Claude-User Allow: / Disallow: /claim/ # Google AI (Gemini training) User-agent: Google-Extended Allow: / Disallow: /claim/ # Perplexity AI User-agent: PerplexityBot Allow: / Disallow: /claim/ # Perplexity live browsing / citation fetcher User-agent: Perplexity-User Allow: / Disallow: /claim/ # Common Crawl (used by many AI models) User-agent: CCBot Allow: / Disallow: /claim/ # Meta FacebookBot (link previews) User-agent: FacebookBot Allow: / Disallow: /claim/ # Cohere AI User-agent: cohere-ai Allow: / Disallow: /claim/ # Amazon / Alexa User-agent: Amazonbot Allow: / Disallow: /claim/ # Apple AI training (Siri, Apple Intelligence) User-agent: Applebot-Extended Allow: / Disallow: /claim/ # =========================================== # AI/LLM Crawlers — Blocked (aggressive training-only) # =========================================== # ByteDance / TikTok (aggressive training crawler) User-agent: Bytespider Disallow: / # Meta AI training (separate from FacebookBot link previews) User-agent: Meta-ExternalAgent Disallow: /