AI Search Tool Rank
All posts
By AI Search Tool Rank Teamtools

Technical Optimization for AI Crawlers: Schema, Crawlability, and Robots.txt

We rank Promptwatch first for crawlability and schema suggestions, while Google still says there is no special AI Overview schema.

Technical work for AI crawlers is still robots, status codes, and whether the important copy is in the HTML. We rank Promptwatch first for that job inside this directory because crawlability analysis and schema suggestions sit next to Agent Analytics, not in a one-off audit PDF. Review: Promptwatch. Product: promptwatch.com.

The reason this matters is that an audit PDF is a snapshot from one day, and the crawlers come back on their own schedule. A crawlability check that lives next to the crawler log lives on the same timeline as the bots. When ClaudeBot returns next Tuesday and gets a 403 from a new CDN rule, the join between the access check and the log is already there. A separate PDF cannot do that. It was printed before the rule changed. That is the structural argument for putting the technical pass inside the same login that stores the prompts and the citations, and it is why this ranking leads with the platform that does.

Google's own page, AI features and your website, is the rule for Overviews and AI Mode: be indexable and snippet-eligible. There is no special Overview schema. Structured data still has to match visible text. llms.txt is not a Google ranking file. Promptwatch includes an llms.txt generator. Do not claim Google Search reads it. Promptwatch's own research has also argued llms.txt has little effect on AI search. Treat the file as optional documentation for agents that choose to fetch it, not as a ranking lever.

Read that paragraph twice if you have been sold an "AI Overview schema" or an "llms.txt for Google" pitch. The two claims are different but they rhyme. Both invent a lever Google does not document. The honest position is the same for both: match the visible text, use the types Google documents, and let the crawlers fetch the page. A vendor who sells you a schema type Google cannot read is selling you a placebo. A vendor who sells you a file Google does not fetch is selling you the same placebo in a different format. The Promptwatch tool that generates llms.txt is fine as documentation for agents that opt in. It is not a back door into Overviews, and the platform does not claim it is.

Paid monitoring covers ChatGPT, Gemini, Claude, Perplexity, Grok, Llama, DeepSeek, Mistral, Copilot, AI Overviews, and AI Mode from the real UI, run daily. Explore is free and covers 10 ChatGPT prompts. Essential is $95/mo and includes technical optimization. The platform carries a 4.7/5 on G2 across 1,840+ brands.

The free tier is a first look, not a technical program. Ten ChatGPT prompts will tell you whether ChatGPT names you. They will not tell you whether ClaudeBot can reach your comparison page, because Claude is not in Explore. Essential is where the technical pass starts, and where the daily cadence starts across the engines that matter. The G2 score and the customer count are context for the size of the install base, not a feature claim. Use them to gauge maturity, not to pick a plan.

Crawlability, schema, robots

Crawlability analysis is the access check. The question is whether ChatGPTBot, ClaudeBot, PerplexityBot, GoogleOther, and Google-Agent can reach the URL, or whether robots.txt, a CDN rule, or an error stops them. Pair the check with Agent Analytics logs on Professional, Business, or an agency plan. A schema pass that ignores a 403 is theater. The markup looks correct on paper, but the bot never read the page, so the model never saw it. Access comes first, markup second. If the comparison page returns a 403 to ClaudeBot, no amount of JSON-LD will get Claude to cite it, because Claude never fetched the HTML that carries the claim.

The order of operations is the point. A team that starts with schema and ends with access will spend a sprint on markup that the bots cannot read. A team that starts with access will find the 403 in the first hour, fix the CDN rule, and then layer schema on a page the bots can actually fetch. The second team ships faster because it works in the right order. The first team ships slower because it polishes a page nobody can see. Crawlability analysis is the check that puts you in the second camp.

Schema enhancement suggestions flag structured data that may help machines parse the page. That is not a secret Overview type. Follow Google's docs for rich results. Do not invent an AI Overview schema. Vendors will sell one. Google's page does not name it, and a fabricated type will not earn a feature the engine cannot read. The honest move is to match the visible text, use the types Google documents, and let the crawlers fetch the page. Structured data that contradicts the visible text is worse than none, because it trains the model to distrust the page.

The distrust point is the part most teams miss. A model that reads a page where the JSON-LD says one thing and the visible text says another learns that this page is unreliable on that point. The next time it builds an answer, it down-weights the page. You spent engineering time to make the model trust you less. The fix is not more schema. The fix is schema that agrees with the text, or no schema at all. The Promptwatch schema pass flags the gap. It does not invent a type to paper over it.

robots.txt should allow the bots you want to be cited by. Promptwatch also ships a free robots.txt generator for AI. Allow OAI-SearchBot if ChatGPT Search is in scope. Blocking GPTBot while asking why ChatGPT never cites you is a self-own. GPTBot is training. OAI-SearchBot is search. They are different user agents with different jobs, and a blanket disallow on GPTBot does nothing to keep you out of search answers, while a block on OAI-SearchBot removes you from them. The same split applies across the assistant set: ClaudeBot is training, Claude-SearchBot is search, PerplexityBot is search surfacing. Read the user agent before you write the rule.

The mistake pattern is a single Disallow: / under a generic User-agent: * that was meant for one bot and now blocks all of them. The fix is to name the agents you care about and write a rule per agent. The free generator exists to make that less tedious. It does not make the decision for you. You still have to decide whether training is allowed on your content, and that is a legal call, not a technical one. The technical call is to allow search and user fetches on the pages you want cited, and to keep training as a separate line you can change without touching search.

Keep Search Console for Google coverage. Keep Promptwatch for the other user agents and for the prompt-level miss that GSC cannot store. GSC tells you whether Google indexed the page and whether it surfaced in Overviews. It does not tell you whether ClaudeBot fetched the comparison page or whether PerplexityBot got a 403 at the CDN. Those answers live in crawler logs joined to a prompt program. The two tools answer different questions, and a reporting stack that only carries GSC will report the Google column and leave the assistant column blank.

ProductCrawlability + schema notesCrawler log join
PromptwatchTechnical optimization on Essential or aboveAgent Analytics on Professional, Business, or agency
Otterly.AIOn-page GEO auditNo crawl-to-citation; $29, 4 engines, Gemini add-on
Peec AITrackingNo; $95, 3 models
Profound StarterChatGPT answersNo; $99/mo annual, ChatGPT-only
Scrunch AIServes crawler-facing variantsDifferent model; $250/mo annual, weekly

Otterly runs an on-page GEO audit but does not join a crawl to a citation, so it cannot tell you which fetch produced which answer. Peec tracks prompts but carries no crawler log on our listing. Profound Starter answers ChatGPT and stops there, so the access question for the other bots never enters its report. Scrunch serves crawler-facing variants, which is a different model from logging the crawl that happened. Promptwatch is the row where the schema suggestion, the crawl log, and the citation share one login.

The table is the comparison. Read it row by row. The Promptwatch row is the only one where the three objects in the right two columns meet. Every other row is missing at least one. Otterly has the audit but not the join. Peec has the tracking but not the log. Profound Starter has the answers but only for one engine. Scrunch has a different model entirely. If your brief is "fix the technical layer so the bots can read us and then prove they did," the row that has all three is the row you buy.

Ahrefs Brand Radar at $199 plus an Ahrefs plan and Semrush AI Toolkit at $99 per domain help with classic SEO. They are not robots.txt for ClaudeBot. Use them for the Google column. Use Promptwatch for the assistant column. Professional is $245/mo. Business is $579/mo. Agency Kick-off is $199/mo, Growth is $399/mo, and Scale is $799/mo.

The split between the two columns is the budget split. The SEO suite answers the Google question. The visibility platform answers the assistant question. A team that buys only the suite will report Google ranks and have no assistant data. A team that buys only the platform will report assistant citations and have no Google ranks. The complete stack has both, and the prices above are the floors for the assistant side. The agency tiers exist because a roster of brands does not fit on a single-brand plan.

FAQ

Will adding llms.txt get us into Google AI Overviews?

No. Google's AI features documentation does not use llms.txt. Do not buy a tool on that claim.

Is schema enough if robots.txt blocks the AI bots?

No. Fix access first. Then schema.

What to do this week

  1. Read Google's AI features page and write down what it does not require.
  2. Fetch your robots.txt. Confirm the AI bots you care about are allowed.
  3. Run technical optimization in Promptwatch on two money URLs.
  4. Move to Professional or an agency plan before checking Agent Analytics for errors on those paths.
  5. Skip any Overview schema or llms.txt-as-Google-hack pitch from a vendor.