A GEO audit — an audit of your Generative Engine Optimization, also called an AI visibility audit — is a manual measurement of how AI engines see, describe and recommend your brand. We run a fixed panel of your buyers’ real prompts through ChatGPT, Claude, Gemini, Perplexity, Grok and DeepSeek, log every answer verbatim with its sources, check what the crawlers behind those engines can actually read on your site, benchmark the competitors the models name instead of you, and hand back a health score out of 100 with a dated 90-day plan.
It is the entry point to everything else we do, and it is complimentary for the requests we take on. You keep the scorecard and the plan whether or not we work together.
why this is a different exam
Your buyer used to open Google, skim ten links and form a view. Now they ask an assistant “who should I pick for X” and get a shortlist of three. If your name is not in that shortlist, the rest of your marketing never gets a hearing — not because you ranked eleventh, but because the model had no reason to name you.
That is a different failure from a ranking failure, and the numbers say so. Across the 100 audits in our published dataset (n = 100, April–July 2026), 93% of brands were known to AI by name and described accurately — and only 4.5% were named when a buyer asked who to pick. Being known and being recommended turned out to be almost unrelated. We found brands with excellent organic traffic and zero presence in AI answers, and brands with almost no traffic that engines quoted first.
So a classic SEO audit will not find this. It measures whether Google can rank you. This measures whether a model can read you, verify you, and quote you.
what you get
| no. | section | what is inside |
|---|---|---|
| 01 | where you stand today | Health Score out of 100 and the numbers behind it: ranked keywords, traffic value, page speed, schema coverage, llms.txt, crawler access. |
| 02 | whether AI names you | Your buyers’ real prompts run live across six engines: who gets named, who gets cited, from which sources, how often you appear — and the prompts where you never do. |
| 03 | who it names instead | Every competitor the models named, benchmarked on the same scale: traffic, citations, entity coverage, their strongest pages, and where they are thin. |
| 04 | what is blocking you | Three critical blockers with severity and the evidence behind each, plus the second-tier list you can hand to a developer. |
| 05 | what to publish | The keyword and prompt map, and the exact pages to ship first, with volumes, difficulty and the intent behind them. |
| 06 | the plan and the targets | Sprint deliverables with exit criteria and a target figure for every KPI we would be judged on. |
| 07 | the hand-offs | Ready-to-apply artefacts, not homework: llms.txt, JSON-LD blocks, redirect and header configs, page briefs. |
how we run it
Ten specialist passes, each one a person looking at a different failure mode, then a synthesis that ranks everything by what it costs you. The order is not cosmetic — each pass depends on what the previous one found.
- Recon. What the domain is, what it was before, and what the open web already says about it. Domain history matters more than people expect: we have seen a previous owner’s reputation still driving the answer a model gives today.
- Technical. Redirects, canonicals, duplicate hosts, security headers, sitemap and robots hygiene — the layer that decides whether anything else is even reachable.
- Crawler reality check. We request your pages as GPTBot, ClaudeBot, PerplexityBot and the rest, and compare what comes back with what a browser sees. This is where client-side rendering quietly deletes brands from the corpus.
- Content and E-E-A-T. Whether there is anything worth quoting, whether a named human stands behind it, and whether the claims can be checked by someone who is not you.
- Schema. The full JSON-LD stack, entity identifiers, and whether the machine-readable version of your company agrees with the human-readable one.
- Entity. Your name across the open web: directories, profiles, knowledge bases, and the collisions where something else owns your name.
- SXO. The page types that win your money queries today, and the gap between those and what you have.
- Market data. Demand by market, the keywords that are genuinely reachable, and the traffic value sitting with each competitor.
- Performance. Core Web Vitals on mobile and desktop, with the render chain that explains them.
- AI visibility panel. The prompt panel, run live, logged verbatim, per engine, per market.
The panel is the part nobody else hands over. Each prompt is a question your buyer actually types, and each answer is stored the way the model produced it, with the sources it leaned on. That panel becomes your baseline. If we work together, every review re-runs the same panel and reports the delta — series, not single runs, because AI answers are volatile enough that one lucky screenshot proves nothing.
what the audits keep finding
The same failures repeat across companies of every size and budget. From the published dataset:
- No entity anchor. Nothing that lets a model confirm you are a real, specific company. Found in close to every audit we ran.
- No machine-readable trust. Licences, jurisdictions and credentials that exist only as images or PDFs, while the negative sources about you are perfectly readable text.
- Anonymous content. No named author, no dates, no sources — the exact profile a model discounts when it has to choose whom to quote.
- Blocked crawlers. Often unintentional, sometimes a managed rule nobody knows about. Blocking GPTBot does not remove you from AI answers; it removes your version of the story and leaves review aggregators and competitors to tell it.
- Nothing shaped like an answer. No comparisons, no pricing, no lists — the page types models quote.
The market average across those 100 audits was 35 out of 100, and 58% of sites scored 35 or below. Not one cleared 60. That is the bad news and the opportunity in the same sentence: almost any category is still winnable from a standing start.
what happens after
You get the report and the 90-day plan, and a call to walk through it. From there one of three things happens, and all three are fine by us: you hand it to your own team, you hand it to another vendor, or we run the plan together as a full-cycle retainer with bi-weekly sprints and a re-measure at every review.
The retainer starts from $3,000/mo (September 2026); the final price is set after the audit. The audit stays complimentary either way. We are referral-first and take a limited number of engagements, so the honest constraint is capacity, not price.
how to read a health score
The score is a weighted roll-up, not a grade curve. It exists so that two audits can be compared and so that a re-audit shows movement, and it is deliberately harsh: nothing in our published dataset of 100 companies cleared 60.
| band | what it usually means | what it costs |
|---|---|---|
| 0–25 | Something structural is broken: crawlers blocked, content rendered client-side only, or the domain carries a reputation that is not yours. | You are absent from the corpus. Nothing downstream can work. |
| 26–40 | The most common band — 58% of the dataset sits at 35 or below. Readable site, no entity, nothing quotable. | Known to AI, never recommended. The 93% / 4.5% gap in one line. |
| 41–60 | Entity exists, some citable content, gaps in consistency and in the sources engines actually read. | Named occasionally, and usually last in the list. |
| 61–80 | Nobody in our sample was here. It would mean consistent identity, citable pages and placement in the sources. | This is what the 90-day plan aims at. |
Two warnings about the number. It is a diagnostic, not a KPI — the KPI is whether engines name you on your own prompt panel. And a high score with zero category recall is possible: a technically immaculate site that nobody cites still loses. We report both, and if they disagree we say which one to believe.
what moving one band actually takes
The score is not a grade to admire; it is a route. Each band has a characteristic bottleneck, and the audit’s 90-day plan is essentially the instruction for crossing into the next one. This is also where our guarantee lives: after the audit we can commit to specific target prompts, scoped against the baseline the audit produced — and the band tells us how bold that commitment can honestly be.
Out of 0–25: engineering weeks, not marketing months. The blockers here are binary — a 403 to AI crawlers, a page that only assembles in the browser, a domain shadowed by someone else’s reputation. Each is a configuration or infrastructure fix with a before/after you can verify the same day. Companies leave this band fastest, because nothing about it is subtle: in our dataset the difference between a blocked site and a readable one was the difference between a 1,133-byte empty shell and the actual product page, on the same URL, in the same second.
Out of 26–40: the entity grind. This is where 58% of the market sits, and where the work is least glamorous — structured data that describes the company rather than decorating pages, named people, an about page a model can verify, identifiers that agree across the open web, an llms.txt that maps the site. None of it produces a visible change to a human visitor, which is exactly why it stays undone. It typically spans two to three sprint cycles, and the re-run usually shows movement on branded probes first: the engines describe you more precisely before they start recommending you.
Out of 41–60: other people’s pages. Above 40 the remaining points are mostly held by third parties — the directories, listicles and communities the engines cite in your category. The baseline already named them, because we logged every source the answers used. Entering them is relationship and publishing work, and it moves at the speed other people publish; this is the band where patience is a strategy rather than an excuse, and where the panel’s per-prompt log matters most, because progress arrives one prompt at a time rather than as a tide.
The bands also explain a pattern clients find counterintuitive: the lower the score, the faster the early gains. A company at 20 can add fifteen points in a month, because its blockers are mechanical. A company at 50 fights for each point, because the remaining ones live on pages it does not own. The audit prices that difference into the plan — and into what we are willing to commit to.
an SEO audit and a GEO audit are not the same document
The overlap is real and we do not pretend otherwise — roughly a third of the technical findings would appear in a good SEO audit too. The rest has no equivalent.
| question | SEO audit | GEO audit |
|---|---|---|
| can it be crawled? | Googlebot, which renders JavaScript | GPTBot, ClaudeBot, PerplexityBot — several of which do not |
| is the markup valid? | Yes, for rich results | Yes, plus whether identifiers connect into one entity graph |
| who are the competitors? | Whoever ranks on the keyword | Whoever the model names in the answer — frequently not the same companies |
| what is the baseline? | Positions and traffic | A logged prompt panel with verbatim answers and cited sources |
| what does success look like? | Higher position, more sessions | Your name inside the answer, and the citation that put it there |
| what is invisible to it? | Everything above | Nothing in the SEO layer — we run it as well, because ranked pages feed citations |
one audit, walked through
Anonymised, from the published set, because it is the clearest example of how the layers stack.
The starting picture. A licensed operator in a regulated category. Real licence, real compliance team, a site that looked fine to its own marketers and scored well in the usual tools.
Pass 3, crawler reality. The site returned 403 to every AI crawler we tested. Not a bug anyone had noticed: the rule lived on the CDN edge, not in the repository, so nobody on the team could see it.
Pass 10, the prompt panel. On branded prompts the engines knew the company. On the question a buyer actually asks — who should I use — it appeared in none of the answers, and two engines volunteered a warning about verifying the brand before paying it. The sources they cited were a review aggregator and a complaint thread, because those were the only readable things about the company.
What the score was made of. Technical passed, content failed on anonymity, entity failed outright, citability was near zero. The blockers were not ranked by severity in the abstract but by cost: unblock the crawlers, publish machine-readable licence and trust data, get named in two independent lists.
What we could not promise. How fast the aggregator would stop being the primary source. We said so in the report, because the honest answer was that it depends on how quickly independent sources appear — and that is partly outside anyone’s control.
what happens if the numbers do not move
Sometimes they do not, and the report is written so that this is visible rather than buried. Three things follow.
First, we show the panel anyway — the same prompts, the same engines, the delta at zero. A flat line is data; a hidden flat line is a sales tactic.
Second, we say which layer failed. Usually it is citability, because it depends on third parties agreeing to publish, and that is the one part of the loop we cannot execute alone.
Third, if two consecutive reviews show nothing, that is a reason to stop rather than to re-sell. We would rather lose a retainer than build a case study out of a client who is not moving.
what the audit is not
- Not a tool report. No dashboard export with a logo on it. Every prompt is run live and logged; every crawler check is a real request.
- Not a keyword list. Keywords are in it, but the panel is prompts — the sentences your buyers type into a chat.
- Not the ceiling. After the audit we can commit to specific target prompts — which answers you take and by when. That guarantee is scoped against your measured baseline, which is exactly why the audit comes first.
- Not a pitch deck. It goes out with whatever number the measurement produced, including when that number flatters a competitor.
the dataset behind it
| what | how much |
|---|---|
| manual audits run | 235+ |
| companies mapped | 3,900+ |
| AI answers analyzed | 93,650+ |
| engines in every audit | 6 — ChatGPT, Claude, Gemini, Perplexity, Grok, DeepSeek |
| audits published in full | 100 — the report |
We work across eight niche groups — web3 & crypto (incl. prediction markets), iGaming, B2B services (development agencies, consulting, legal, accounting and tax, recruiting, marketing agencies, corporate services), fintech & payments, AI / AI tools, B2B SaaS, e-commerce & DTC, and real estate & construction — so in most categories we already know what the AI shelf looks like before your audit starts.
One caveat we state everywhere: the public report is not the plan. It names the patterns we found across a hundred companies — roughly 20% of the work. The other 80% is your own numbers, and it only exists once someone measures your brand. That is what this audit is.
what an audit looks like when it is finished
Two audits from the dataset are public in full, with the client's written approval, because the shortest way to judge a report is to read one. Both are published the way we deliver them: baseline first, every figure dated, every run logged, including the runs where nothing moved.
| sample | the baseline | the latest logged run |
|---|---|---|
| Manimama, EU crypto-law firm | 22 June 2026: health score 45 of 100, named in 0 of 20 buyer prompts across four engines, 1,004 blog posts and none in Google US top-3. | August 2026: named in 20 of 20 prompts, first in ChatGPT on two of them, second in Europe on Perplexity, cited from its own service page. |
| INC4, AI and blockchain studio | 14 May 2026: health score 27 of 100, robots.txt and sitemap returning 404, 177 of 549 backlinks pointing at dead pages, Google's AI Overview resolving the name to an IEEE conference. | 2 July 2026: health score 53; by 28 July all four engines naming it organically; on the frozen AI-for-fintech panel 8 of 52 answers naming it and first place on the main question by 31 August 2026. |
Read either one the way you would read your own: the seven sections above are the same, the prompt panel is frozen on day one and never edited, and the timeline shows what shipped next to what the engines did about it. If the format is what you want for your domain, the request takes a website and three priority geos, and the audit is complimentary.