# robots.txt — zxiaohan.com # # 立场:本站**欢迎会引用来源的 AI / 答案引擎爬虫**,内容就是写来被发现、 # 被引用、被生成式搜索召回的。见 /llms.txt 的结构化站点摘要。 # # 中文搜索引擎(百度 / 搜狗 / 360 / 神马)与字节、华为的爬虫同样放行—— # 本站有中文内容且公司业务面向国内。 # # 但欢迎不等于门户大开。只索引不回流的商业 SEO 工具、数据倒卖爬虫、 # 以及不守规矩的高频抓取器,一律拒绝——它们消耗带宽,不带来任何可见度。 # # robots.txt 语义提醒:爬虫只遵守**最匹配的那一组**规则。所以下面被点名的 # 爬虫都不受末尾 `User-agent: *` 的 Crawl-delay 约束。 # ══ 一、AI / 生成式答案引擎(明确放行)════════════════════════════ # 这些会在回答里给出处链接,是本站 GEO 策略的核心受众。 # OpenAI User-agent: GPTBot Allow: / User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / # Anthropic (Claude) User-agent: ClaudeBot Allow: / User-agent: Claude-Web Allow: / User-agent: Claude-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: anthropic-ai Allow: / # Perplexity User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # Google (Gemini / AI Overviews grounding) User-agent: Google-Extended Allow: / # Apple Intelligence User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / # Microsoft / Copilot User-agent: bingbot Allow: / # Common Crawl(喂给大量开源模型,抓取频率温和) User-agent: CCBot Allow: / # 其他答案引擎 User-agent: Amazonbot Allow: / User-agent: meta-externalagent Allow: / User-agent: FacebookBot Allow: / User-agent: DuckAssistBot Allow: / User-agent: YouBot Allow: / User-agent: cohere-ai Allow: / User-agent: MistralAI-User Allow: / # 字节(豆包 / 抖音搜索)——抓取量凶,但国内可见度靠它。 User-agent: Bytespider Allow: / # ══ 二、传统搜索引擎(明确放行,不受 Crawl-delay 限制)════════════ User-agent: Googlebot Allow: / User-agent: Googlebot-Image Allow: / User-agent: Slurp Allow: / # ── 中文 / 中国市场搜索引擎 ──────────────────────────────────────── # 本站有中文内容(/rbtx/zh/)且公司业务面向国内,这些必须显式放行—— # 不点名的话它们会落进末尾的 `User-agent: *`,被 Crawl-delay: 10 限速。 # 百度 User-agent: Baiduspider Allow: / User-agent: Baiduspider-render Allow: / User-agent: Baiduspider-image Allow: / # 搜狗 User-agent: Sogou web spider Allow: / User-agent: Sogou inst spider Allow: / # 360 User-agent: 360Spider Allow: / User-agent: HaosouSpider Allow: / # 神马(UC / 阿里) User-agent: Yisouspider Allow: / # 华为 Petal 搜索 User-agent: PetalBot Allow: / # ══ 三、拒绝:商业 SEO 工具爬虫 ═══════════════════════════════════ # 抓全站建自己的外链数据库然后卖订阅,对本站零回流。 User-agent: AhrefsBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: DotBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: BLEXBot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: serpstatbot Disallow: / User-agent: Barkrowler Disallow: / User-agent: SeekportBot Disallow: / User-agent: ZoominfoBot Disallow: / # ══ 四、拒绝:高频抓取 / 数据倒卖,且无引用回流 ═══════════════════ # 注:Bytespider(字节 / 豆包)与 PetalBot(华为 Petal 搜索)曾在此段被拒, # 2026-08-26 已分别移到第一、二段放行——本站面向国内的中文内容和公司业务需要它们的 # 可见度。Bytespider 抓取量确实凶,靠 Cache Rule 和 WAF 限速去控量, # 而不是靠封禁换取安静。 User-agent: Diffbot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Timpibot Disallow: / User-agent: omgili Disallow: / User-agent: omgilibot Disallow: / User-agent: magpie-crawler Disallow: / User-agent: Scrapy Disallow: / # ══ 五、其余一切 ══════════════════════════════════════════════════ # 允许,但限速。上面点名放行的爬虫不受这条约束。 User-agent: * Allow: / Crawl-delay: 10 Disallow: /quant/app/data/ # ══ Sitemap ══════════════════════════════════════════════════════ # 索引文件,下挂 core / aptamer / quant 三张分区表,共 1000+ URL。 # 每个 URL 都带准确的 lastmod;日期归档页标 changefreq=yearly, # 明确告诉爬虫这些页发布后不再变化,不必反复回抓。 Sitemap: https://zxiaohan.com/sitemap.xml # 百度明确不处理 sitemap 索引文件(站长平台原话:「请勿提交索引型 sitemap」), # 360、神马、头条也未必读。三张分区表单独再列一遍;Google / Bing 重复读到无害。 Sitemap: https://zxiaohan.com/sitemap-core.xml Sitemap: https://zxiaohan.com/sitemap-aptamer.xml Sitemap: https://zxiaohan.com/sitemap-quant.xml