Every audit we hand over opens with one number out of 100. It is the fastest way to say how much of your visibility problem is machinery rather than marketing — and it is worthless if you cannot see how it was produced. So this page publishes the method: the categories, the weights, the bands, and the limits of what the number can tell you.
The score answers a narrow question: if an AI engine tried to read, verify and quote this site today, how far would it get? It is not a ranking prediction and not a measure of how often models currently name you — that is a separate measurement, described further down.
seven categories, one weighted sum.
Each category is scored 0–100 on its own evidence, then multiplied by its weight. The weighted results add up to the final score, so the scale needs no further normalisation.
| category | weight | what it scores |
|---|---|---|
| content quality | 23% | E-E-A-T signals, readability, freshness, author attribution, thin or duplicated pages. |
| technical SEO | 22% | HTTP headers, security, canonicals, mobile, JS rendering, hreflang — whether a crawler gets a page at all. |
| on-page / SXO | 20% | Page-type match against real queries, structure, internal linking, the path a buyer takes once they land. |
| schema | 10% | JSON-LD across five to seven page types: Organization, Product or Service, FAQPage, Person, BreadcrumbList. |
| performance | 10% | Core Web Vitals and delivery — measured, or inferred and marked as inferred when field data is missing. |
| GEO / AI search | 10% | The AI-readiness block, itself weighted — see the breakdown below. |
| images | 5% | Formats, weight, alt text, whether product and team imagery is machine-describable. |
Two documented variants exist, and when we use one the audit says so on the cover: a GEO-weighted run lifts GEO to 15% and drops performance to 5% for brands whose whole problem is AI-side; and since May 2026 the GEO block can be split into two halves — 5% site readiness and 5% live AI-citation measurement.
five dimensions of being quotable.
The GEO category is scored the same way — five dimensions, each 0–100, weighted into one number.
| dimension | weight | what it scores |
|---|---|---|
| citability | 25% | Whether the page states claims a model can lift verbatim: specifics, numbers, comparisons, prices, named sources. |
| structural readability | 20% | Headings, lists, tables, answer-shaped blocks — the difference between text a model can quote and prose it must summarise. |
| authority & brand signals | 20% | Entity records, third-party mentions, Wikipedia and directory status, and whether your name collides with a bigger entity. |
| technical accessibility | 20% | AI crawler access, llms.txt, robots rules, server-side rendering — whether GPTBot and friends are let in at all. |
| multi-modal content | 15% | Images, video and data that carry meaning without the surrounding page. |
five bands.
| score | tier | what it means in practice |
|---|---|---|
| 0–39 | critical | Foundation issues. The engine is missing, misreading or refusing the site; content work would be spent on ground that cannot hold it. |
| 40–59 | poor | Several categories failing at once. Usually readable, rarely verifiable, almost never quotable. |
| 60–74 | fair | Solid base, clear opportunities. This is where authority work starts paying instead of leaking. |
| 75–89 | good | Refinement and authority building — the machinery is no longer the constraint. |
| 90–100 | excellent | Maintenance and edge cases. |
what the market actually scores.
Across the hundred audits published in our report, the average is 35 out of 100. More than half of audited businesses — 58% — score at or below that, and not one site cleared 60. The distribution: 8% land in 0–20, 50% in 21–35, 34% in 36–50, 8% in 51–60.
By niche the picture splits: web3 and crypto average 37, iGaming 24 — the worst category we measure — and agencies and B2B land in a 28–45 band. The practical reading is not that everyone is bad at this. It is that the shelf is still cheap: a score in the sixties is currently enough to be the site an engine reaches for.
what actually adds points.
The engineering-heavy fixes are the ones that move the number, and they move it in chunks rather than percentages. A full JSON-LD stack — Organization, Product or Service, FAQPage, Person, BreadcrumbList — is typically one of the two largest single lifts at +15 to +25 points, because the schema category starts near zero in most audits: the typical score we record is 0–4 out of 100.
The rest, in the order we usually sequence them: let the AI crawlers in and render server-side; publish an llms.txt; give the business a machine-readable identity (an /about and /team that state who you are); then convert prose into answer-shaped blocks that can be quoted without rewriting. Content and authority work comes after — not because it matters less, but because it compounds on a base that holds.
what the score is not.
It is not a Lighthouse score. Performance is one tenth of it. A perfect 100 in PageSpeed changes this number by a few points at most.
It is not a ranking or traffic forecast. It measures readiness, not demand. A site can score 70 in a category nobody searches for.
It is not our AI-visibility score. That is a separate measurement of how often engines actually name you — run on a live prompt panel, with its own scale and its own bands, where 0–15 already counts as critical. Readiness and visibility move together eventually, but never on the same day: fixing the machinery is what makes the citations possible, not what buys them.
It is not permanent. We re-measure the same panel and the same categories every fourteen days, because engines change their minds and so do the sources they read.
how the inputs are gathered.
Ten specialist passes feed the seven categories: a full crawl of every route with headers and rendering checks, JSON-LD detection across five to seven page types, an AI-crawler and llms.txt check, a SERP-backwards pass on the queries that actually matter, keyword clustering and competitor gaps, backlink and authority signals, performance, and two demand passes through DataForSEO. On top of that a prompt panel built for your category is run live through six engines, with every answer logged next to the sources it cited.
Every figure carries its date and locale, and anything inferred rather than measured is labelled as inferred. Where a dimension has too little data to score honestly — backlinks on a brand-new domain, for instance — we suppress the number instead of guessing it.
questions about the score.
what is a good GEO health score?
Sixty is the practical line: it is where the machinery stops being the constraint, and in a hundred audits not one site reached it. Below 40 a brand is effectively invisible to the models — not absent from the web, but absent from the answer. Between 40 and 59 you are usually readable and rarely quotable.
can I calculate it myself?
The model is on this page, so yes in principle: score each of the seven categories 0–100 on evidence you can show, apply the weights, and add them up. The hard part is not the arithmetic, it is the evidence — particularly the GEO block, which needs live answers from several engines and the sources they cited, not an opinion about your content.
is this the same as an SEO audit score?
It overlaps and then diverges. Classic SEO categories carry 75% of the weight, because the pages engines quote are usually pages that already rank — but the GEO block, the citability dimension and the live prompt panel measure something SEO tools do not: whether a model can lift a claim from your page and stand behind it.
how fast does the number move?
The technical and schema work lands within one measurement cycle — fourteen days — and typically shows up as +15 to +25 points. Authority and entity signals take months, because they depend on third parties updating what they say about you.
do you publish the score for your own site?
We publish the method, the dataset and the distribution behind it in the 100-audit report. For an individual site the score is only meaningful next to its category rivals measured the same day, which is what the audit does.
The number is the opening line of the audit, not the product. What follows it is the part that matters: which specific answers you are missing, who is being named instead, and what to change first.