[{"data":1,"prerenderedAt":138},["ShallowReactive",2],{"product-/en/product/ai-crawlability-en":3},{"id":4,"title":5,"body":6,"description":6,"extension":7,"meta":8,"navigation":94,"path":132,"seo":133,"stem":136,"__hash__":137},"content_en/1.product/ai-crawlability.yml","Ai Crawlability",null,"yml",{"category":9,"hero":10,"takeaway":20,"features":24,"technical":44,"benefits":72,"comparison":88,"faq":106,"cta":122},"Features",{"eyebrow":11,"headline":12,"description":13,"links":14},"Get Found, the first rung","Check which AI crawlers [your site refuses]{class='text-primary'}","A page an AI crawler cannot fetch will never be cited. Most brands fail this rung without knowing: a robots.txt written for Googlebot still refuses GPTBot.",[15],{"label":16,"to":17,"size":18,"color":19},"14-day free trial","https://app.trustdata.tech","lg","primary",{"label":21,"text":22,"attribution":23},"Key takeaway","When a citation arrives with no matching crawl, it came from **training data**, not retrieval. Rewriting the page will not move it. Blocking training while allowing retrieval is a coherent position; blocking all three with one inherited wildcard is what usually happens.","A citation with no crawl record cannot be fixed by editing the page",{"title":25,"description":26,"items":27},"Four questions your log files can answer and your rank tracker cannot","Detection happens server-side in first-party tracking, so this is observed traffic rather than a third-party estimate.",[28,32,36,40],{"title":29,"description":30,"icon":31},"Which engine actually fetched the page","Name the engine behind every visit. Each known AI crawler user-agent maps to what it feeds, OpenAI, Claude, Gemini, Perplexity, Mistral, DeepSeek, Meta AI and the rest, and a visit you cannot attribute is labelled untagged rather than quietly dropped.","i-lucide-radar",{"title":33,"description":34,"icon":35},"Whether it came to train or to retrieve","Decide what to block on purpose rather than by accident. Each bot carries a purpose, training corpus, search index, on-demand fetch or mixed, and blocking training while allowing retrieval is a coherent strategy that almost nobody is actually running.","i-lucide-split",{"title":37,"description":38,"icon":39},"Which of your pages are turned away","See which money pages your own robots.txt refuses. Half or more of them blocking AI bots caps the Get Found rung at weak however heavily the rest of the site is crawled, because a blocked page cannot be cited.","i-lucide-shield-off",{"title":41,"description":42,"icon":43},"How long a crawl takes to become a citation","Stop guessing how long to wait. Per page and per engine you get the gap between first crawl and first citation over **90 days**, so you know whether a publishing change has had time to land.","i-lucide-timer",{"title":45,"description":46,"features":47},"The four outcomes of a crawl","Crawl records and citation records are joined on a full outer join, so the cases where one side is missing survive rather than disappearing. Each is a different problem.",[48,52,56,60,64,68],{"title":49,"description":50,"icon":51},"Crawled and cited","Both sides present, so the latency between them is measurable. This is the working case, and the distribution of that latency per engine is what tells you how fast your category moves.","i-lucide-circle-check",{"title":53,"description":54,"icon":55},"Crawled, not cited","The engine fetched the page and has not cited it in **90 days**. Retrieval is not the problem. This is an extractability or authority problem, which is rungs 3 and 4, and the page audit is where you go next.","i-lucide-file-x",{"title":57,"description":58,"icon":59},"Cited with no crawl record","The engine cites the page but never fetched it inside your history. That usually means the citation comes from training data rather than live retrieval, which is a very different asset: it does not update when you edit the page.","i-lucide-brain",{"title":61,"description":62,"icon":63},"Insufficient history","Fewer than **14 days** since the first event, so no conclusion is drawn. It reads as insufficient history rather than as a zero, because a page crawled yesterday has not failed to be cited.","i-lucide-hourglass",{"title":65,"description":66,"icon":67},"Per-bot visit counts","First seen, last seen and visit count per bot per page over **90 days**, so a crawler that stopped coming is visible as a change rather than as an absence you have to notice.","i-lucide-activity",{"title":69,"description":70,"icon":71},"Engine-level rollups","Crawl volume by engine over time, so you can see one engine going quiet while the others hold steady, which is usually a robots or a rate-limit problem rather than a content one.","i-lucide-chart-no-axes-column",{"title":73,"description":74,"items":75},"What you do with it","Crawl data is only useful if it changes a decision. These are the three it changes most often.",[76,80,84],{"title":77,"description":78,"icon":79},"Fix robots.txt with evidence","Turn the robots argument into a list. Most AI bot blocks are inherited rather than chosen, a wildcard written for scrapers or a plugin default, and seeing which named bots are refused on which pages ends the debate.","i-lucide-file-cog",{"title":81,"description":82,"icon":83},"Separate a retrieval problem from a content problem","Find out which of three different weeks of work you are in. Crawled but not cited is a content problem, never crawled is an access problem, and cited with no crawl record is a training-data asset you cannot edit your way into or out of.","i-lucide-git-branch",{"title":85,"description":86,"icon":87},"Know when to stop waiting","Know when a change has genuinely failed. The observed crawl-to-citation latency for your own pages and engines tells you how long absence has to last before it means anything, measured rather than assumed.","i-lucide-timer-reset",{"title":89,"description":90,"items":91},"Against the usual ways of checking this","Most teams either read raw logs by hand or assume it is fine.",[92,97,99,101,104],{"feature":93,"trustdata":94,"ga4":95,"other":96},"AI crawler visits identified server-side",true,false,"Log tools, unclassified",{"feature":98,"trustdata":94,"ga4":95,"other":95},"Bot mapped to the engine it feeds",{"feature":100,"trustdata":94,"ga4":95,"other":95},"Training against retrieval against on-demand",{"feature":102,"trustdata":94,"ga4":95,"other":103},"robots.txt checked against your money pages","Manual",{"feature":105,"trustdata":94,"ga4":95,"other":95},"Crawl-to-citation latency",[107,110,113,116,119],{"label":108,"content":109,"defaultOpen":94},"How does TrustData see AI crawler visits?","Detection happens server-side in the first-party tracking layer, which inspects the user agent of every request before any JavaScript runs. AI crawlers do not execute JavaScript, so a client-side analytics tool cannot see them at all. That is why crawler activity is missing from most dashboards rather than showing as zero: the measurement never had a chance to happen.",{"label":111,"content":112},"Why does the purpose of a crawler matter?","Because the decisions are different. A training crawler is building a corpus that will be baked into a future model. A search-index crawler is maintaining an index that answers questions today. An on-demand fetch happens because a user asked something right now and the engine went to read your page for that answer. Blocking training while allowing retrieval is a defensible position on your content rights. Blocking all three because a wildcard rule caught them is what usually happens, and it costs you citations for no considered reason.",{"label":114,"content":115},"What does it mean when a page is cited but was never crawled?","It generally means the citation comes from the model's training data rather than from live retrieval. That is worth knowing because it behaves differently: editing the page will not update what the model says, and the citation can vanish with a model refresh you do not control. It is also why probes run in two modes, with live web search and without. A brand strong without search and weak with it exists in the model but is not being retrieved; the reverse means the opposite.",{"label":117,"content":118},"How is the crawl-to-citation latency measured?","Crawl events and citation records are joined per page and per engine across 90 days of history, on a full outer join so that pages with only one side present are kept rather than dropped. Where both sides exist, the gap between the first crawl and the first citation is the latency. Where only one exists, the row is classified: crawled but not cited, or cited with no crawl record. Under 14 days since the first event nothing is concluded, and the row reads insufficient history.",{"label":120,"content":121},"Does blocking AI crawlers hurt my visibility?","On the retrieval side, directly and completely: an engine that cannot fetch a page cannot cite it, which is why the Get Found rung caps at weak when half or more of your audited money pages block AI bots. On the training side it is a genuine trade rather than a mistake, and some brands make it deliberately. TrustData's job is to make sure it is a decision. The common case we see is neither: a rule written years ago for a different problem, still quietly refusing the crawlers that matter now.",{"title":123,"description":124,"links":125},"Find out who is actually reading your site","14-day free trial. Crawler monitoring is part of AI Visibility, included in every plan.",[126,128],{"label":16,"to":17,"size":127,"color":19},"xl",{"label":129,"to":130,"variant":131,"size":127},"See the full readiness ladder","/en/product/ai-visibility","outline","/product/ai-crawlability",{"title":134,"description":135},"AI crawler access monitoring, from your logs","See which AI crawlers reach your pages, whether they came to train or to retrieve, and how long it takes a crawl to turn into a citation.","1.product/ai-crawlability","TYkgzY0Zm8F9aYgJu-OR73Srub5crmZz4kxlaaanV5E",1786740929936]