To be picked up as a source in answers from ChatGPT, Claude, Perplexity and Google, your site has to be crawlable first, readable next, and quotable last. PiScan measures all three across five layers.
126 sites averaged 80/100. Scores come from PiScan’s own measurement; you can run the same scan on your own address.
Brand names and icons are used only to identify which site each measurement belongs to. We have no commercial relationship with these brands, nor their endorsement.
SCANNING
WHY IT'S DIFFERENT
AI visibility isn't measured with a checklist. The only way to learn whether a crawler reaches your page is to send a request carrying its identity. Five things set PiScan apart:
Most tools read robots.txt, say "GPTBot allowed" and stop there. We fetch the same page separately as GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Googlebot, then compare the responses. Servers that welcome Google while refusing AI identities surface right here. We also state the limit plainly: this test shows identity-sensitive behaviour, it doesn’t prove real crawler access.
The page is split at heading boundaries the way a real answer-engine pipeline would, and every chunk is inspected individually. We don’t say "your page is good" — we say "5 of your 12 chunks open without context".
Not how many bytes your page weighs, but how many tokens it occupies in a model’s context — and what share of that weight is genuinely content.
We compare the FAQ questions declared in your structured data against the page body. Markup that doesn’t match isn’t merely useless — it can breach structured-data policies and cost you rich-result eligibility.
When we see a block, we first separate "is this AI-specific or a general bot wall?". If Googlebot is refused too, we don’t report it as a critical finding — because it isn’t one.
MEASUREMENT FRAMEWORK
The layers are deliberately ordered: if a crawler can't get in, text quality is moot; if text can't be extracted, its structure is moot. The weights follow the same logic.
Can the crawler actually get in?
robots.txt is parsed separately for 16 AI crawlers; then the same page is requested carrying those crawlers’ identities. That is how we catch the case where the rule file says "come in" while the server looks at the identity and shuts the door.
What does a crawler without JavaScript see?
Most AI crawlers don’t execute JS. We take the raw HTML, strip the script and style scaffolding, and measure the real text left behind. The result: how many KB of your page is content, and how many tokens the page occupies in a model’s context.
Does it survive being split into chunks?
An answer engine doesn’t quote your page — it quotes your CHUNK. So we perform the same split and ask each chunk two questions: does it open without context, and does it name the subject? A chunk that opens with "This reduces costs" tells a reader almost nothing on its own.
Does your structured data build a graph?
Not "is there JSON-LD" but: are the nodes linked by @id, is there an anchor to external authorities via sameAs, and do the declared FAQ questions actually appear on the page — and a mismatch can breach Google’s structured-data policies, which may cost you rich-result eligibility.
Is there a clean passage to quote?
We measure headings that mirror real user questions, the short self-contained answer beneath them, verifiable figures, freshness declarations and boilerplate overlap with internal pages. When the same boilerplate dominates every page, retrieval systems struggle to tell one page from another.
HOW IT WORKS
Give a real content page rather than the homepage; the measurement becomes far more meaningful.
About 10 requests go out: the page, robots.txt, llms.txt, the sitemap, 5 crawler identities and 2 sampled internal pages. Total time is capped at 55 seconds.
A separate score for each of the five layers, the crawler parity matrix, and a "what to do" note beside every finding.
If you want to go through the report with a specialist, you can get in touch in one click. Entirely optional.
MONTHLY SCAN
AI crawler behaviour isn't fixed: robots rules change, CDN settings get updated, new pages appear. A one-off scan shows you that day; a monthly scan shows you the change.
FAQ
It measures how accessible, readable and quotable your site is for AI crawlers. There are five layers: access (can a crawler get in), extractability (how much text survives without JavaScript), chunkability (do retrieval chunks stay meaningful), the semantic layer (does structured data build an entity graph) and citability (can an answer engine find a passage to quote).
No. Classic SEO tools look at ranking signals; PiScan measures readiness for answer engines (ChatGPT, Claude, Perplexity, Google AI Overviews). The robots.txt, sitemap and llms.txt checks are only the entrance — the real measurements are the crawler parity test, raw-HTML efficiency and chunk health.
Three free scans a day, resetting each night at midnight (Europe/Istanbul). Your allowance is tracked server-side: refreshing the page, opening a new tab or closing the browser and coming back does not reset it. Rescanning the same address within 10 minutes serves a cached result and uses no allowance. If a scan fails for technical reasons, the allowance is refunded.
Two things: a salted hash of your IP address (the raw IP is never written anywhere) and a random identifier cookie (pis_uid). Both may count as personal data under Türkiye’s Personal Data Protection Law (KVKK) and, where applicable, the GDPR; they are processed solely to count daily scans and prevent abuse, never for advertising or profiling, never shared with third parties, and deleted after three days.
No. PiScan only reads: it downloads your page and rule files like an ordinary visitor. No form is submitted and no data is sent. Roughly 10 requests are made per scan, and the same address can only be scanned a few times per minute.
Because we measure the page a crawler sees, not the page you see. PiScan doesn’t run JavaScript — just like GPTBot, ClaudeBot and PerplexityBot. If the gap is large, that gap is precisely the point.
It is done by imitating the crawler identity (User-Agent). Real crawlers additionally pass IP verification, so where we see a block the real crawler may get through. PiScan states this openly in the report: if Googlebot is refused as well, it does not present that as "AI is blocked" but marks it as a general bot wall.
Start with the critical findings; those are the items that directly cut visibility. Every finding carries a "what to do" note. If you like, we can go through the results together and turn them into a prioritised plan.
CONTACT
ASK AI ABOUT DIJITALPI
Opens your chosen assistant with a ready research prompt. It reads the site live and answers.
We use cookies to improve your experience and to analyse site traffic anonymously. Details: Cookie Policy · Privacy