CCBot
Common Crawl ・ AI学習 ・ 取得日 2026-09-10
非営利の Common Crawl がオープンなWebクロールデータセットを作るためのクローラー。生成されたデータセットは研究や各種AIの学習に広く使われる。
運営元
Common Crawl
用途
AI学習
robots.txt の挙動
robots.txtに従う
当社robots.txt
未記載
robots.txt の挙動
公式は User-agent: CCBot / Disallow: / を robots.txt に書くことをブロック方法として案内。CCBotを騙る偽装クローラーの報告があるため、公式IPレンジの逆引き(crawl.commoncrawl.org)での検証を推奨。
User-agent 文字列(公式)
CCBot/2.0 (https://commoncrawl.org/faq/)CCBot をブロックする robots.txt
User-agent: CCBot
Disallow: /出典
https://commoncrawl.org/ccbot (Common Crawl 公式(CCBot) ・ 取得日 2026-09-10)