AI Crawler Activity in Website Logs: Cloudflare AI Bots, Visibility Tools, Google AI Overviews Crawler, and Bing Copilot Crawler Analytics (2026)
There is no AI Overviews crawler and no Copilot crawler in your logs. We map each AI surface to the bot that feeds it, then rank the tools that read that activity.
Search your server logs for an "AI Overviews crawler" and you will come up empty. It does not exist. Google's crawler documentation (updated 2026-07-14) says Googlebot crawl preferences apply to "Google Search (including Discover and all Google Search features)." AI Overviews and AI Mode are Search features, so Googlebot is the crawler that feeds them. Copilot works the same way on the Microsoft side. Bing's webmaster guidelines say Bing and Copilot search experiences "rely on the same core crawling, indexing, and ranking foundation as traditional search." The Copilot crawler in your logs is Bingbot.
That collapses two of the four names in the title into bots you have been logging for years. The other two, Cloudflare's AI bot reporting and the visibility tools built on top of logs, are where the new data lives. This page maps the surfaces to the user agents, explains what Cloudflare shows without extra setup, and ranks the tools we would use to read AI crawler activity in 2026.
Which bot feeds which AI surface
| AI surface | What appears in your logs | The control that matters |
|---|---|---|
| Google AI Overviews and AI Mode | Googlebot | Your Googlebot rules. No separate user agent |
| Gemini Apps and Vertex AI grounding, Gemini training | Nothing. Google-Extended is a robots.txt token | Google-Extended in robots.txt |
| Generic Google fetches | GoogleOther | Google says it doesn't affect any specific product |
| Bing Copilot | Bingbot | robots.txt for crawl; NOINDEX, NOARCHIVE, NOCACHE for use |
| ChatGPT Search | OAI-SearchBot | Its own robots.txt rule |
| OpenAI model training | GPTBot | A separate rule from OAI-SearchBot |
| A ChatGPT user opening your page | ChatGPT-User | User-initiated, not the Search control |
The Google-Extended row trips up a lot of teams. Google says it "doesn't have a separate HTTP request user agent string," so you will never see it in a log line. It is a robots.txt token for Gemini training and for grounding in Gemini Apps and Vertex AI. Google also says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal." Blocking Google-Extended does not take you out of AI Overviews. If someone on your team blocked it expecting that, the logs will show Googlebot crawling as usual, because nothing changed for Search.
On Bing, the guidelines list blocking Bingbot in robots.txt as a mistake. They also separate crawl from use. robots.txt controls crawling, not indexing. NOINDEX keeps a URL out of Bing, Copilot and grounding results. NOARCHIVE "prevents content from being used in Copilot responses and grounding results." NOCACHE limits Copilot to the URL, title and snippet. So a page can be crawled by Bingbot every day and still never be quoted in Copilot, and the log alone will not tell you why. Check the meta tags too.
One more caution before you trust any of these rows. User agents can be spoofed. Google tells you to verify Googlebot by reverse DNS (the host should resolve to googlebot.com) or against its published IP ranges. Microsoft offers a Verify Bingbot tool. A burst of "Googlebot" traffic from an unknown host is not Google, and it should not count toward your AI Overviews story.
What Cloudflare shows without extra tooling
If your site sits behind Cloudflare, AI Crawl Control (formerly AI Audit) is already in your dashboard on every plan. It has Overview, Crawlers and Metrics tabs. You get requests by crawler and by operator (OpenAI, Microsoft, Google, ByteDance, Anthropic, Meta), allowed versus unsuccessful requests, robots.txt violations, status codes, paths and hosts. You can allow or block each crawler from the same screen, and pay per crawl is in private beta.
The free plan detects crawlers by user-agent string only, and the Metrics tab covers the past 24 hours. Referral counts are on paid plans. Enterprise with Bot Management adds detection IDs, configurable timeframes and pay per crawl. The data is also exposed through the GraphQL Analytics API if you want to pull it elsewhere.
What Cloudflare does not do is connect a crawl to anything downstream. It will tell you GPTBot fetched /pricing and got a 200. It will not tell you whether ChatGPT cited /pricing in an answer, for which prompt, or whether anyone clicked through and bought.
Ranked: tools for reading AI crawler activity
1. Promptwatch. Agent Analytics logs AI crawlers in real time (Promptwatch names ChatGPTBot, ClaudeBot, PerplexityBot, GoogleOther and the Meta AI crawler), tracks errors, and draws the crawl-to-citation path from a fetched page to the answers that cite it. It ingests from Cloudflare (Enterprise via Logpush, other plans through an auto-deployed Worker, with the record orange-cloud proxied), AWS CloudFront, Fastly, Vercel, Netlify, Akamai, Google Cloud CDN and custom HTTP. For the Google and Bing surfaces, where the crawler is the ordinary search bot, the part that adds information is answer monitoring. Promptwatch checks the real interfaces of Google AI Overviews, AI Mode and Copilot alongside ChatGPT, Perplexity and the others, and citation analytics shows which of your pages each answer used. Visitor analytics then counts the AI-referred sessions and conversions. Crawler logs start on Professional at $245/mo (25M logs) and Business at $579/mo (100M). Essential at $95/mo has no crawler-log allowance listed. Agencies get 10M logs on Kick-off at $199/mo.
2. Cloudflare AI Crawl Control. The best free starting point if you already proxy through Cloudflare. Good for spotting blocks and 4xx spikes by operator. The 24-hour window on free plans makes trend work hard.
3. Ahrefs Bot Analytics. Server-side bot analytics, free during beta, with 12 bot categories and an "AI bots" toggle. Logs arrive through Cloudflare Logpush or a Worker, a Vercel log drain, Fastly or an Amazon data stream. It is separate from Ahrefs Brand Radar, which is the prompt side.
4. Semrush Log File Analyzer. Upload or connect server logs and filter for GPTBot, ChatGPT-User, OAI-SearchBot and ClaudeBot. Useful for an audit, less so for daily monitoring.
5. Raw logs and a script. Free, complete and slow. You will spend most of the effort on reverse DNS verification.
We compare the crawler tools head to head in AI crawler visibility tool for bot activity on your website. For setup, see how to connect Cloudflare crawler logs to Promptwatch.
Our recommendation
Read crawler activity in two halves. For OpenAI, Anthropic, Perplexity and Meta, the bot is AI-specific, so the log itself is the signal. For AI Overviews and Copilot, the bot is Googlebot or Bingbot, so the log mostly confirms access and the real question is whether the answer cited you. A tool that only reads logs gives you the first half. That is why we rank Promptwatch first: on Professional it puts crawler logs, the Google and Copilot answers, citations and AI-referred conversions in one project. Start on Promptwatch, then compare configured prices across the full rankings.