Get Found, the first rung

Check which AI crawlers your site refuses

A page an AI crawler cannot fetch will never be cited. Most brands fail this rung without knowing: a robots.txt written for Googlebot still refuses GPTBot.

Key takeaway

When a citation arrives with no matching crawl, it came from training data, not retrieval. Rewriting the page will not move it. Blocking training while allowing retrieval is a coherent position; blocking all three with one inherited wildcard is what usually happens.
A citation with no crawl record cannot be fixed by editing the page

Four questions your log files can answer and your rank tracker cannot

Detection happens server-side in first-party tracking, so this is observed traffic rather than a third-party estimate.

Which engine actually fetched the page

Name the engine behind every visit. Each known AI crawler user-agent maps to what it feeds, OpenAI, Claude, Gemini, Perplexity, Mistral, DeepSeek, Meta AI and the rest, and a visit you cannot attribute is labelled untagged rather than quietly dropped.

Whether it came to train or to retrieve

Decide what to block on purpose rather than by accident. Each bot carries a purpose, training corpus, search index, on-demand fetch or mixed, and blocking training while allowing retrieval is a coherent strategy that almost nobody is actually running.

Which of your pages are turned away

See which money pages your own robots.txt refuses. Half or more of them blocking AI bots caps the Get Found rung at weak however heavily the rest of the site is crawled, because a blocked page cannot be cited.

How long a crawl takes to become a citation

Stop guessing how long to wait. Per page and per engine you get the gap between first crawl and first citation over 90 days, so you know whether a publishing change has had time to land.

The four outcomes of a crawl

Crawl records and citation records are joined on a full outer join, so the cases where one side is missing survive rather than disappearing. Each is a different problem.

Crawled and cited

Both sides present, so the latency between them is measurable. This is the working case, and the distribution of that latency per engine is what tells you how fast your category moves.

Crawled, not cited

The engine fetched the page and has not cited it in 90 days. Retrieval is not the problem. This is an extractability or authority problem, which is rungs 3 and 4, and the page audit is where you go next.

Cited with no crawl record

The engine cites the page but never fetched it inside your history. That usually means the citation comes from training data rather than live retrieval, which is a very different asset: it does not update when you edit the page.

Insufficient history

Fewer than 14 days since the first event, so no conclusion is drawn. It reads as insufficient history rather than as a zero, because a page crawled yesterday has not failed to be cited.

Per-bot visit counts

First seen, last seen and visit count per bot per page over 90 days, so a crawler that stopped coming is visible as a change rather than as an absence you have to notice.

Engine-level rollups

Crawl volume by engine over time, so you can see one engine going quiet while the others hold steady, which is usually a robots or a rate-limit problem rather than a content one.

What you do with it

Crawl data is only useful if it changes a decision. These are the three it changes most often.

Fix robots.txt with evidence

Turn the robots argument into a list. Most AI bot blocks are inherited rather than chosen, a wildcard written for scrapers or a plugin default, and seeing which named bots are refused on which pages ends the debate.

Separate a retrieval problem from a content problem

Find out which of three different weeks of work you are in. Crawled but not cited is a content problem, never crawled is an access problem, and cited with no crawl record is a training-data asset you cannot edit your way into or out of.

Know when to stop waiting

Know when a change has genuinely failed. The observed crawl-to-citation latency for your own pages and engines tells you how long absence has to last before it means anything, measured rather than assumed.

Against the usual ways of checking this

Most teams either read raw logs by hand or assume it is fine.

AI crawler visits identified server-side

TrustData

GA4

Other tools

Log tools, unclassified

Bot mapped to the engine it feeds

TrustData

GA4

Other tools

Training against retrieval against on-demand

TrustData

GA4

Other tools

robots.txt checked against your money pages

TrustData

GA4

Other tools

Manual

Crawl-to-citation latency

TrustData

GA4

Other tools

FAQ

Direct answers.

14-day free trial

Find out who is actually reading your site

14-day free trial. Crawler monitoring is part of AI Visibility, included in every plan.