เครื่องมือสร้าง robots.txt สำหรับ AI bot
เลือกนโยบายสำหรับ AI bot แต่ละประเภท ตั้งค่าทับสำหรับ crawler รายตัว แล้วคัดลอกไฟล์ที่เป็นไปตาม RFC 9309 token สำหรับการฝึกโมเดล การค้นหา และการดึงข้อมูลตามคำขอของผู้ใช้แยกจากกัน คุณจึงปฏิเสธประเภทหนึ่งและอนุญาตอีกประเภทได้
# บล็อก AI crawler ไม่ให้เข้าถึงทั้งเว็บไซต์
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: GoogleOther
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: Amazonbot
User-agent: MistralAI-Training
User-agent: CCBot
User-agent: Bytespider
User-agent: Timpibot
Disallow: /
# อนุญาต AI crawler ยกเว้น path ส่วนตัว
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: Googlebot
User-agent: Google-CloudVertexBot
User-agent: Google-Agent
User-agent: Google-GeminiNotebook
User-agent: Gemini-Deep-Research
User-agent: GoogleAgent-Mariner
User-agent: Applebot
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: meta-webindexer
User-agent: meta-externalfetcher
User-agent: facebookexternalhit
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: DuckAssistBot
User-agent: MistralAI-User
User-agent: MistralAI-Index
User-agent: cohere-ai
User-agent: YouBot
User-agent: Diffbot
Disallow: /admin/
Disallow: /api/
# crawler อื่นทั้งหมด
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Disallow: /admin/
Disallow: /api/
Sitemap: https://example.com/sitemap.xml
robots.txt เป็นข้อกำหนดโดยสมัครใจ ChatGPT-User, Perplexity-User และ fetcher อื่น ๆ ที่ทำงานตามคำขอของผู้ใช้ระบุว่ากฎเหล่านี้อาจไม่มีผลกับตนเอง และ agent บางตัวก็ไม่มี token เลย หากต้องการบังคับใช้การบล็อก ให้ตรวจสอบคำขอเทียบกับรายการ IP หรือ signature ที่เผยแพร่ไว้ แล้วปฏิเสธคำขอที่ firewall หรือ WAF
ไฟล์ที่สร้างขึ้นทำงานอย่างไร
crawler ทำตามกลุ่มของตัวเอง
เมื่อ crawler พบ token ของตัวเองในบรรทัด User-agent ก็จะทำตามกฎในกลุ่มนั้น และข้ามกฎใต้ User-agent: * นี่จึงเป็นเหตุผลที่ generator ใส่ path ส่วนตัวซ้ำสำหรับบอตที่คุณอนุญาตด้วยชื่อ และกำหนด Disallow: / เพียงรายการเดียวให้กับบอตที่ถูกบล็อก
บาง token ไม่ได้ใช้ crawl
Google-Extended และ Applebot-Extended ไม่มี user agent ของตัวเอง Googlebot และ Applebot เป็นผู้ดึงข้อมูล ส่วน extended token จะบอก Google และ Apple ว่าสามารถนำเนื้อหาไปฝึกโมเดลได้หรือไม่ นอกจากนี้ Google ยังใช้ Google-Extended กับ Gemini grounding ด้วย
ผู้ให้บริการส่วนใหญ่แยกการฝึกโมเดลกับการค้นหา
OpenAI, Anthropic, Amazon และ Mistral ต่างมี token แยกสำหรับการฝึกโมเดลและการค้นหา ดังนั้น preset No training จะยังให้คุณอยู่ในผลิตภัณฑ์ค้นหาของผู้ให้บริการเหล่านี้ ส่วน Meta เป็นข้อยกเว้น: meta-externalagent ครอบคลุมทั้งการฝึกโมเดลและการทำดัชนี
Applebot ใช้กฎของ Googlebot หากไม่มีกลุ่มของตัวเอง
Apple ระบุว่าหากไฟล์ของคุณไม่มีกลุ่ม Applebot, Applebot จะทำตามกฎของ Googlebot แทน นอกจากนี้ Applebot และ Amazonbot จะไม่สนใจ Crawl-delay ขณะที่ Anthropic ระบุว่า ClaudeBot ปฏิบัติตามกฎนี้
คำถามเกี่ยวกับ robots.txt และ AI bot
จะบล็อก GPTBot แต่ยังให้ค้นหาใน ChatGPT ได้อย่างไร
กำหนด Disallow ให้ GPTBot และอนุญาต OAI-SearchBot ไว้ GPTBot เก็บข้อมูลสำหรับฝึกโมเดล, OAI-SearchBot ทำดัชนีหน้าเว็บสำหรับการค้นหาใน ChatGPT และ ChatGPT-User จัดการการดึงข้อมูลตามคำขอของผู้ใช้ OpenAI อธิบายทั้งสามรายการไว้ในหน้าเดียว พร้อมรายการ IP แยกสำหรับแต่ละรายการ
การบล็อก Google-Extended จะทำให้เว็บไซต์หายจาก Google Search หรือไม่
Google-Extended เป็น control token ที่ไม่มี user agent ของตัวเอง โดย Googlebot เป็นผู้ crawl การกำหนด Disallow ให้ token นี้จะบอก Google ไม่ให้นำเนื้อหาของคุณไปใช้ฝึกโมเดล Gemini หรือทำ grounding ส่วนกฎของ Googlebot จะเป็นตัวกำหนดว่าเนื้อหาใดจะถูก crawl เพื่อใช้ใน Search
AI crawler เคารพ Crawl-delay หรือไม่
บางตัวเคารพ Anthropic ระบุว่า ClaudeBot ปฏิบัติตามกฎนี้ ส่วน Apple และ Amazon ระบุว่า Applebot และ Amazonbot ไม่ปฏิบัติตาม การจำกัดอัตราคำขอที่เซิร์ฟเวอร์เป็นวิธีที่เชื่อถือได้กว่า
ทำไม ChatGPT-User ยังเข้ามาเยี่ยมชมหลังจากบล็อกแล้ว
OpenAI ระบุว่ากฎใน robots.txt อาจไม่มีผลกับ ChatGPT-User เนื่องจากมีผู้ใช้ร้องขอหน้าเว็บ Perplexity-User, Amzn-User และ meta-externalfetcher ก็มีข้อยกเว้นคล้ายกัน หากผู้ให้บริการเผยแพร่รายการ IP ไว้ ให้ตรวจสอบคำขอเทียบกับรายการดังกล่าว แล้วบล็อกที่ firewall
Content-Signal ใน robots.txt ใช้ทำอะไร
ระบุว่าเนื้อหาที่ดึงมาแล้วนำไปใช้ได้อย่างไร ผ่านสัญญาณ 3 รายการ ได้แก่ search, ai-input และ ai-train ซึ่งตั้งค่าแต่ละรายการเป็น yes หรือ no สัญญาณเหล่านี้แสดงเพียงความต้องการและไม่ได้บล็อกสิ่งใด Cloudflare's Markdown for Agents จะเพิ่ม ai-train=yes, search=yes และ ai-input=yes ให้โดยค่าเริ่มต้น ดังนั้นหากใช้ฟีเจอร์นี้ ให้ตรวจสอบผลลัพธ์ที่ได้