A GEO audit checks whether AI engines can read your site, verify that you are a real and specific company, and find something on it worth quoting when a buyer asks who to pick. This is the full list of checks behind our manual audit: 40 of them, grouped by the ten passes the audit runs in. Where the published dataset of 100 audits states a benchmark for a check, it is next to the check, dated; where it does not, there is no number.

The order is not cosmetic: each pass depends on what the previous one found, and a failure early in the list makes everything after it moot. The benchmarks are grim on purpose: the average health score across the 100 audits was 35 out of 100, 58% of sites scored 35 or below, and not one cleared 60 (the report, published 17 July 2026). The engines knew 93% of the brands by name and named 4.5% in category answers; the checklist is the anatomy of that gap.

01 · recon

What the domain is, what it was before, and what the open web already says about it.

  1. Domain history. Previous owners, brands and reputations. A model can still be answering with the last owner's story; ours was a betting brand on similar domains, hence the disambiguation section in our llms.txt.
  2. What the open web says. Search the brand with “legit” and “scam”; that is the corpus an engine reads when a buyer asks whether you are trustworthy.
  3. Name collisions. Everything else that shares your name. In the INC4 case Google's AI Overview resolved the firm's name to an IEEE conference and a UN committee first (snapshot 27 July 2026, the case).

02 · technical

The layer that decides whether anything else is reachable at all.

  1. One canonical host. www and non-www, http and https, trailing slashes: one of each resolves, the rest redirect once.
  2. Redirect chains and dead URLs. Every historical URL lands on content. In the INC4 audit 177 of 549 backlinks pointed at dead pages (baseline 14 May 2026); 41 redirect rules closed them.
  3. robots.txt and sitemap.xml return 200. In the same audit both returned 404, as did two pages in the main menu.
  4. Canonicals, duplicates and hreflang. One URL per page, one language per URL, and the tags that say so.
  5. Security headers and cache policy. The boring set; a marketer never sees a missing header, a trust heuristic does.

03 · crawler reality check

Requesting your pages the way the engines do, and reading what comes back. Benchmark: 50% of audited sites were unreadable to AI crawlers through client-side rendering, blocked bots or outright 403s (n = 100, April to July 2026).

  1. Fetch as each AI user-agent. GPTBot, ClaudeBot, PerplexityBot and the rest, on home, services, pricing and about. Log the status and the actual bytes of text served.
  2. Compare with what a browser sees. One site in the dataset scored 100/100 for SEO in Lighthouse while the same URL returned 1,133 bytes of empty shell to GPTBot. Both were true at the same moment.
  3. Find where the block lives. Often a CDN rule nobody can see from the repository; one licensed operator in the dataset served 403 to every AI crawler from the edge.
  4. robots.txt names the AI crawlers explicitly. Allow or disallow is a decision, not a default; blocking removes your version of the story, not the story.

04 · content and E-E-A-T

Whether there is anything worth quoting, and whether a named human stands behind it. Benchmark: 89% of audited sites published anonymous content, with no about or team pages and no article authors (n = 100).

  1. A named author on every article, with credentials that can be checked outside your site.
  2. Dates and sources on claims. An undated, unsourced number is the profile a model discounts when choosing whom to quote.
  3. /about and /team exist, return 200, and say who you are. Founding year, legal entity, where you operate from.
  4. Numbers agree across pages. Client counts, years in business, team size: one fact set. Three versions of you means a model trusts none of them.
  5. Trust signals are text, not images. A licence that exists only as a badge or a PDF does not exist for the model; the complaint thread about you is perfectly readable.

05 · schema

The machine-readable version of your company, and whether it agrees with the human-readable one. Benchmark: 94% of audited sites had no JSON-LD at all; the typical schema score we record is 0 to 4 out of 100, and a full stack is typically one of the two largest single lifts to a health score, +15 to +25 points (how the score works).

  1. JSON-LD is present and validates. Broken FAQ markup and a generic Organization with no business type count as absent.
  2. The full stack. Organization, WebSite, Service or Product, Person, FAQPage, BreadcrumbList, and Article on the posts.
  3. One organisation node with a stable identifier, referenced by everything else rather than re-described on every page. Disconnected markup reads as two weak entities instead of one strong one.
  4. Markup matches the page. FAQ answers identical to the visible text, ratings only where reviews exist, sameAs pointing at your own profiles. Contradicting markup is worse than none.

06 · entity

Your name across the open web, and whether a model can confirm you are one specific company. Benchmark: close to 100% of audited sites had no entity anchor; not one had a Wikidata record (n = 100).

  1. A Wikidata item, and for mature brands a Wikipedia article. These are the sources models verify entities against.
  2. One fact set on Crunchbase, LinkedIn and the directories. Same name, founding year and description; a Crunchbase entry describing last year's product becomes last year's story told as fact.
  3. Google Knowledge Graph and AI Overview resolve the name to you. Ask for the company by name and check every fact in the answer.
  4. Location and identifiers agree. Directory listings with a different city each get repeated by engines as offices. Ours did; the correction lives in llms.txt and on the about page.

07 · SXO

The page types that win your money queries today, and the gap between those and what you have. On-page and SXO carries 20% of the health score; citability is the largest dimension of the GEO block at 25%.

  1. Page-type match. For each money query, the page type the engines quote (comparison, pricing, list, definition), and whether you have one.
  2. Answer-shaped blocks. A definition near the top that can be lifted in one paragraph, a table, a price signal, an FAQ whose markup matches the text word for word.
  3. Structural readability. Headings, lists and tables rather than prose a model has to summarise; 20% of the GEO block.
  4. Internal links that describe the structure. Hubs link to spokes in body copy, not only in the footer; siblings link to each other.

08 · market data

Demand by market, and where the value currently sits.

  1. Demand per market and language. Which queries and prompts are real in each geo you sell into; the audit covers one to five languages per brand (methodology, the report).
  2. Competitor traffic value. What each rival's organic visibility would cost as ads; $62 a month next to a competitor's $3,342 is donating the category (example from the report).
  3. The two competitor sets. Who ranks for the keyword, and who the models name in the answer. Frequently different lists; the second one matters here.

09 · performance

Core Web Vitals and the render chain that explains them; 10% of the health score, and never the reason a brand is missing from an answer on its own.

  1. Core Web Vitals on mobile and desktop, from field data where it exists, and marked as inferred where it does not.
  2. The render chain. What blocks first paint, and whether the content buyers read exists only after JavaScript runs.
  3. Page weight as a crawler sees it. A page a crawler abandons before the content is a readability failure, not a speed one.

10 · AI visibility panel

The part no tool report hands over: your buyers' real prompts, run live, logged verbatim, per engine, per market. Benchmark: 93% brand recall, 4.5% category recall; 87 of the 100 audited brands were known to the engines by name and recommended someone else.

  1. A fixed panel of branded probes (“tell me about X”, “is X legit”) and category prompts (“who should we pick for Y”), worded once and never edited; 2 to 14 live prompts per brand in the published audits.
  2. Six engines: ChatGPT, Claude, Gemini, Perplexity, Grok and DeepSeek, each measured separately, because they answer differently.
  3. Every answer logged verbatim with its sources. Who was named, in what position, from which source, and whether the source was yours.
  4. Tone. Whether an engine describes you neutrally, hedges, or warns the buyer to verify you before paying. In iGaming, 2 of 3 audited brands were described as high-risk or a possible scam (n = 100 dataset, iGaming subset).
  5. A series, not a run. The same panel re-run on a schedule and reported as a delta against the baseline; one screenshot proves nothing.

the benchmarks in one table

100 GEO audits, april to july 2026 · 009 Agency, published 17 july 2026
metricvaluechecklist section
average health score35 / 100all
scoring 35 or lower58%all
scoring above 600%all
score distribution8% at 0 to 20 · 50% at 21 to 35 · 34% at 36 to 50 · 8% at 51 to 60all
no entity anchor~100%06 entity
no JSON-LD94%05 schema
anonymous content89%04 content
llms.txt missing or 40487%03 crawlers
unreadable to AI crawlers50%03 crawlers
brand recall93%10 panel
category recall4.5%10 panel
known, never recommended87 of 10010 panel
web3 and crypto: average score37 / 100 · llms.txt present ~20% · Wikidata 0 of audited brandsniche
iGaming: average score24 / 100 · category recall 0 of 10 prompts · crawlers blocked 2 of 3 brandsniche
agencies and B2B: score range28 to 45 · category recall 0 of 13 promptsniche

Download the CSV: 100-geo-audits-2026.csv (also as JSON). The file is the fixed study, n = 100, and its numbers are not restated as the sample grows; the wider dataset behind the method now covers 160+ manual audits, 2,800+ companies and 67,200+ AI answers (August 2026). Reuse is fine with attribution to 009 Agency, per the licence.

how to read your own result

Count the checks you fail in sections 03, 05 and 06 first. They are the mechanical half of GEO, they move fastest, and a site that passes them is already ahead of roughly nine in ten audited companies. Then look at section 10 on its own: a clean site with zero category recall is the common case, and it means the remaining work is on pages you do not own, the directories, roundups and lists the engines cite. That is slower, and no checklist shortens it.

If you would rather have the 40 run properly, that is the audit: manual, five to ten working days, complimentary for the requests we take on, and yours to keep whether or not we work together. Vocabulary is in the glossary; the technical checks are expanded in technical GEO.