User-agent: * Allow: / Disallow: /admin/ Sitemap: https://www.surveycake.com/sitemap.xml # ── AI 搜尋引擎:允許(會把流量帶回來,是 GEO 曝光來源)── User-agent: GPTBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / User-agent: ClaudeBot Allow: / User-agent: Claude-User Allow: / User-agent: anthropic-ai Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / User-agent: Google-Extended Allow: / User-agent: Applebot-Extended Allow: / # ── 純訓練用途、不帶流量:限速 ── User-agent: CCBot Crawl-delay: 10 User-agent: meta-externalagent Crawl-delay: 10 User-agent: Amazonbot Crawl-delay: 10 # ── 高耗流量且不帶回訪:封鎖 ── User-agent: Bytespider Disallow: / # ── 2026-08-18 追加:搜尋引擎限速 ── # 背景:本帳單週期(8/06–9/05)13 天已用 63.82 GB,推估月底 135–152 GB。 # 證據:max-age 一年的核心 JS/CSS 被請求 219,439 次(主檔 71,884 次、每日 5,530 次 # 「首次載入」),17K 的 favicon 被抓 48,948 次。這些檔案正常訪客只會抓一次, # 如此高的次數只可能來自不帶快取的客戶端=爬蟲。 # 取捨:AI 搜尋引擎(GPTBot/ClaudeBot/PerplexityBot 等)維持 Allow 不限速, # 因為它們會把流量帶回來,是 GEO 曝光來源。這裡只對「純索引、不帶回訪」的 # 一般搜尋爬蟲限速。 # ⚠️ Google 官方明示 Googlebot 忽略 Crawl-delay,要限速必須到 Search Console # 設定「檢索頻率」;Bingbot 與 Yandex 則會遵守。 User-agent: bingbot Crawl-delay: 5 User-agent: Yandex Crawl-delay: 10 User-agent: Baiduspider Crawl-delay: 10 User-agent: SemrushBot Disallow: / User-agent: AhrefsBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: DataForSeoBot Disallow: /