CrawlPodScan your site

GPTBot

GPTBot is OpenAI's web crawler for collecting content used to train its generative AI foundation models. Per OpenAI's own documentation, it identifies itself with the User-Agent GPTBot/1.4 and respects robots.txt — disallowing it tells OpenAI a site's content shouldn't be used for training.

Not the only OpenAI bot

OpenAI operates separate crawlers for separate purposes, and they don't all behave the same way: OAI-SearchBot surfaces sites in ChatGPT's search features (blocking it removes a site from ChatGPT search answers, distinct from training), OAI-AdsBot checks ad page safety, and ChatGPT-User fetches pages a user directly asks ChatGPT to visit — OpenAI's own documentation notes robots.txt rules may not apply to that last one, since it's a user-triggered action rather than a background crawl.

On Shopify, WordPress, and Next.js

Allowing GPTBot means explicitly permitting it in robots.txt — some hosting defaults or security plugins block unfamiliar user agents by default. Check with a free scan.

AI crawler · ClaudeBot · PerplexityBot