Methodology

Everything that sits behind the number

21 checks, three pillars weighted differently, and one rule that stops an average from hiding a fatal problem. Here it is in full, so you can argue with it.

The process

Four steps, none of them with surprises

  1. 1

    We request, and we say who we are

    Your page, your robots.txt and your sitemap, with the agent ParsigoBot. We don't pretend to be GPTBot in order to check whether GPTBot can get in: we evaluate your robots.txt with the same rules they apply.
  2. 2

    We don't run JavaScript

    Just like AI crawlers. What we analyse is the HTML exactly as it leaves your server, which is precisely what they receive.
  3. 3

    We walk through several pages

    Up to 12, chosen from your sitemap and the links on the home page, spread across different sections. The home page tends to be the least representative page of a site.
  4. 4

    We evaluate with deterministic rules

    There are no language models in the analysis. The same conditions always give the same result, today and in six months' time.

21 checks

What we check, exactly

This list comes straight out of the analysis engine, it isn't written by hand: if we add a check, it shows up here on its own. The weight is what it is worth within its pillar.

Access

Can they get in?45 % of the score
  • Access for the bots that write answers

    blocks

    The agents that search live and cite with a link. Blocking them costs you real visibility.

    14 pts
  • Page indexability

    blocks

    A noindex in the tag or in the HTTP header leaves the page out, even when the crawler gets in.

    12 pts
  • Server response

    blocks

    If the page does not answer with a 2xx, there is nothing to read.

    10 pts
  • Anti-bot challenges

    blocks

    A badly tuned WAF hands a challenge to legitimate agents. Your robots.txt says yes and your infrastructure says no.

    8 pts
  • robots.txt availability

    blocks

    A server error on robots.txt is read as a blanket ban on crawling.

    6 pts
  • Secure connection (HTTPS)

    Basic requirement. Without it, several agents don't even try.

    4 pts
  • Redirect chain

    Every hop is a chance to lose the crawler along the way.

    4 pts
  • Response time (TTFB)

    An agent resolving a live question is in a hurry: someone is waiting on the answer.

    4 pts
  • Training bots (for information)

    informational

    Blocking them is a legitimate editorial decision and does not affect whether you get cited. We tell you; it costs you no points.

Readability

Can they read you?35 % of the score
  • Content available without JavaScript

    blocks

    No relevant AI crawler runs JavaScript. If the content is assembled in the browser, what they see is an empty page.

    20 pts
  • Fallback without JavaScript (noscript)

    A patch, not a fix: it is there to warn, not to replace.

    5 pts
  • Main heading (H1)

    The most basic signal about what the page is about.

    5 pts
  • Declared language

    With no declared language, the content can end up in the wrong group.

    4 pts
  • Main content marked up

    Marking up with <main> or <article> separates the content from the menu and the footer.

    4 pts

Structure

Can they understand you?20 % of the score
  • Structured data (JSON-LD)

    Saying outright what otherwise has to be inferred from the prose.

    10 pts
  • JSON-LD validity

    One comma too many and the whole block is discarded. There is no graceful degradation.

    6 pts
  • Canonical URL

    If the same page is reachable through several URLs, the signal is split between them.

    6 pts
  • Declared entity types

    Declaring something is not the same as declaring who you are: that takes an entity type.

    5 pts
  • Page title

    The first thing anyone reads of the page, on it and away from it.

    5 pts
  • XML sitemap

    The list of what you want found, without depending on links.

    5 pts
  • Meta description

    The summary shown when somebody finds you.

    3 pts

The distinction people get wrong most often

Training is not citation

Blocking GPTBot keeps your content out of a model's training, but it doesn't stop ChatGPT from citing you: that takes OAI-SearchBot and ChatGPT-User. Same with Google-Extended, which only governs Gemini's training and not the AI Overviews. That is why we check the 19 agents separately and only penalise the blocking of those that affect citation.

They answer right now

6 agentspenalised if blocked
They go out to read when someone asks a question, and they cite with a link. Blocking them costs you visibility: it is the only thing we penalise.
  • ChatGPT-UserOpenAI
  • Claude-UserAnthropic
  • Perplexity-UserPerplexity
  • DuckAssistBotDuckDuckGo
  • MistralAI-UserMistral AI
  • AmazonbotAmazon

They feed the search index

6 agentspenalised if blocked
They build the index the answers are drawn from. Blocking them has a cost too.
  • OAI-SearchBotOpenAI
  • Claude-SearchBotAnthropic
  • PerplexityBotPerplexity
  • GooglebotGoogle
  • BingbotMicrosoft
  • ApplebotApple

They collect for training

7 agents
Their content doesn't come back as a citation. Blocking them is a legitimate editorial decision, and it doesn't lower your score.
  • GPTBotOpenAI
  • ClaudeBotAnthropic
  • Google-ExtendedGoogle
  • Applebot-ExtendedApple
  • CCBotCommon Crawl
  • meta-externalagentMeta
  • BytespiderByteDance

The score

How it's worked out, and why it isn't an average

The pillars are not worth the same

If they can't get in, nothing else gets to matter.

  • Access45 %
  • Readability35 %
  • Structure20 %

What we can't see doesn't pass

If your server doesn't return the HTML, we don't invent a verdict about something we haven't seen: the check drops out of the calculation and is declared separately. Rounding it up would be lying with a number.

One blocker sets a ceiling

6 of the 21 checks take no points away: if they fail, the overall score cannot go above 39. An average would hide one fatal failure behind fifteen minor wins.

What each band means

  • Ready

    90100

  • Ready with reservations

    7089

  • Problems found

    4069

  • Blocked

    039

The five possible outcomes

  • Pass
    We checked it and it's fine. It scores all its points.
  • Warning
    It works, but not as it should. It scores half.
  • Fail
    We checked it and it's wrong. It scores nothing.
  • Inconclusive
    We couldn't check it. It drops out of the calculation: it counts neither as a pass nor as a failure.
  • Not applicable
    The check makes no sense in your case. It also drops out of the calculation.

What we don't measure

We don't check whether an assistant cites you today

The available research suggests that not even the most closely correlated external signal explains more than a small fraction of why a model recommends a brand, and that most of the behaviour cannot be observed from outside. Promising that prediction would mean selling something nobody can stand behind. Parsigo stops at the prerequisite, which is measurable and stable over time.

The other three things we don't do

Now you know how it's worked out. Go and see your score

The 21 checks against your domain, in a few seconds. Free and with no sign-up.

Analyse my site