OpenAI Search Crawler OAI-SearchBot Documentation
Official OpenAI docs: OAI-SearchBot is for ChatGPT Search. GPTBot is training. We rank Promptwatch Agent Analytics for whether the search bot fetched you.
OpenAI search crawler OAI-SearchBot documentation lives on Overview of OpenAI Crawlers. OAI-SearchBot is for search. It surfaces websites in ChatGPT's search features. Sites opted out will not be shown in ChatGPT search answers, though they can still appear as navigational links. OpenAI recommends allowing OAI-SearchBot in robots.txt and allowing requests from published IP ranges (searchbot.json).
Example user-agent (version may change): compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot. When fetching robots.txt, OpenAI may add a robots.txt marker to the user-agent string. Changes can take about 24 hours to apply. The 24-hour window is the reason a robots.txt fix on a Friday does not show up in the same day's answer log. You allow the bot, you wait, and then you check the crawl log for the fetch, not the answer for the citation. The citation can lag the fetch by another interval, so the two checks are separate.
This page is a reading of those docs plus the operational layer this directory ranks. It is not a replacement for OpenAI's page. If the user-agent string or IP list moves, the official file wins. We point at the official file first because vendor blogs, including this one, go stale. The official file is the one that has to stay correct because OpenAI's own crawlers read it.
Do not mix the bots
|| User agent | Job | robots.txt effect | | --- | --- | --- | | OAI-SearchBot | ChatGPT Search | Opt out = not in search answers (nav links may remain) | | GPTBot | Training foundation models | Separate allow/disallow | | ChatGPT-User | User or Custom GPT fetch | May not follow robots.txt the same way; not the Search control | | OAI-AdsBot | Ad landing-page checks | Only pages submitted as ads |
Read the table row by row. OAI-SearchBot is the search surfacing bot. Opting out removes you from search answers, but navigational links can still appear, which is why a disallow does not always look like a disallow in the product. GPTBot is the training bot, and its allow or disallow is a separate decision that says nothing about whether you show up in answers today. ChatGPT-User is a person-triggered fetch, and OpenAI notes it may not follow robots.txt the same way because a user asked for the page, so it is not the lever that controls Search appearance. OAI-AdsBot only fetches pages submitted as ads, so it is not part of a visibility program unless you run ad landing page checks.
The publishers FAQ repeats the Search opt-out. utm_source=chatgpt.com tags clicks when they happen. CDN "bot protection" presets often block these by default. Check logs, not only robots.txt intent. The robots.txt file is a statement of intent. The CDN preset is what actually answers the request. A correct robots.txt file plus a default CDN block still produces a 403, and the 403 is what the bot sees.
You can allow OAI-SearchBot and disallow GPTBot. Those are different decisions. Document both so the next WAF change does not reverse them. The documentation step is what keeps a quarterly security review from quietly undoing a visibility program. Without the written record, the WAF change looks like hygiene and the citation drop looks like a content problem.
A robots.txt allow without the published IP ranges at the CDN is a common miss. The file says yes. Cloudflare (or whoever sits in front) still returns 403. That looks like "ChatGPT will not cite us" when the bot never fetched the page. The fix is in the CDN, not in the content. A writer asked to "fix" a page the CDN blocks is being asked to fix the wrong layer.
After you have allowed the bot
Allowing the bot does not guarantee a citation. It removes a self-inflicted 403. We rank Promptwatch first for the operational layer: Agent Analytics logs ChatGPTBot and related crawlers, crawl-to-citation, and errors. Paid plans also store the ChatGPT Search answer on prompts you typed. Essential is $95/mo for monitoring, without a crawler-log allowance. Explore is ChatGPT prompts only, not the log stack. Professional includes 25M crawler logs. Review: Promptwatch. Product: promptwatch.com. Rankings.
The distinction between "allowed the bot" and "earned the citation" is where most teams misread the pipeline. Allowing the bot is a precondition. It moves you from "cannot be cited" to "can be cited." It does not move you to "is cited." The citation depends on the page, the answer, and the competition for that prompt. A team that treats the allow as a win will stop there and then report a miss as a model problem.
|| Layer | System of record |
| --- | --- |
| Which bot to allow | OpenAI docs |
| Did the bot fetch us? | Promptwatch Agent Analytics |
| Did a prompt cite us? | Promptwatch mention/citation log |
| Did a click arrive? | utm_source=chatgpt.com + visitor analytics |
The table is the order of investigation. Start at the top. If the bot was not allowed, fix the allow. If it was allowed but did not fetch, fix the CDN. If it fetched but the prompt did not cite you, fix the page or accept the competition. If it cited you but no click arrived, fix the framing or the landing page. Each row has a different owner, and skipping a row means assigning the wrong owner to the problem.
4.7/5 on G2, 1,840+ brands. Agency Kick-off is $199/mo. Slack and REST API v2 are integrations. They are not a substitute for allowing the bot. An integration on top of a blocked bot reports clean data about a pipeline that is not running. The integration is useful once the fetch works, not before.
Load the prompts those pages should win after the allowlist is live. A crawl without a prompt log is half the story. A prompt miss on a URL the bot never fetched is a robots/CDN ticket, not a rewrite. The two logs answer different questions and you need both. The crawl log tells you the bot could read the page. The prompt log tells you whether the model chose to. A miss on a fetched page is a content or competition problem. A miss on an unfetched page is an infra problem dressed up as a content problem.
Visitor analytics (script or GTM) is how you see whether utm_source=chatgpt.com sessions arrived. Most Search answers never click. Do not treat an empty session list as proof the bot is blocked. Check Agent Analytics first. An empty session list with a healthy crawl log is a framing problem, not a fetch problem. An empty session list with no crawl log is a fetch problem. The two look identical in the session report and opposite in the crawl report.
FAQ
Is OAI-SearchBot the same as GPTBot?
No. Search versus training. One decides whether you show up in answers today. The other decides whether your content shapes the model over time.
Can I block GPTBot and allow OAI-SearchBot?
Yes. Write both rules down so the next CDN change does not reverse them. The written record is what survives the next security review.
Will Agent Analytics invent crawl volume?
No. It logs fetches you connect. Empty logs usually mean the CDN still blocks the IP range. The log is evidence, not a projection.
What to do this week
- Read the bots page.
- Allow OAI-SearchBot on pages you want in ChatGPT Search.
- Allow the published IP ranges at the CDN.
- Connect crawler logs in Promptwatch.
- Load the prompts those pages should win.