The GEO category is about a year old. Nobody in it has a long track record, most claims cannot be verified from outside, and the vocabulary is easy to copy. So the useful filter is not reputation but whether an agency can show you the measurement it is selling.

Five questions do most of the work. Ask them of every agency you shortlist, including us.

1. what prompts would you measure us on?

A real answer is a list of sentences your buyers would actually type, split into branded and category. A vague answer — “we track AI visibility” — means there is no panel, and without a panel there is no baseline and no delta. This one question eliminates most of a shortlist.

2. which engines, and how often?

Engines disagree with each other constantly; measuring one and generalising is not measurement. Ask for the count and the cadence, and ask whether they store the answers verbatim. Answers are volatile enough that a single run proves very little, so a series is the only honest reporting format.

3. show me the last delta you reported to a client

Redacted is fine. What you are checking is whether the artefact exists at all, and whether it includes runs where nothing moved. An agency that only has improvements to show is either very new or only showing you half the data.

4. what do we receive, exactly?

Ask for the file list. In a healthy engagement it looks like: llms.txt, schema blocks, server config, crawler rules, page briefs, and the logged panel. If the answer is “a strategy document”, ask who implements it — the gap between recommendation and implementation is where most retainers quietly die.

5. what happens if the numbers do not move?

The answer tells you whether you are buying a process or a promise. Ours: we show the flat line, we say which layer failed, and two consecutive empty reviews are a reason to stop rather than to re-sell. Any answer that cannot conceive of failure is a sales script.

three red flags

  • Placement guarantees before any measurement. A commitment to target prompts is only honest after a baseline exists. A guarantee sold before anyone has measured anything tells you how the rest of the engagement will be run.
  • Bought mentions sold as authority. In one of our audits the only readable corpus about a brand was its own paid promotion; the engines flagged the sources as promotional and issued a high-risk verdict. Volume without independent confirmation makes answers worse.
  • No published method. Nine of ten agencies in our own comparison state no methodology and no engine count. In a field selling measurement, that gap is the whole story.

on pricing, since nobody says it

Almost no agency in this category publishes a number, which means buyers cannot self-qualify and every conversation starts with a sales call. We publish ours to break that: the manual audit is complimentary, the retainer starts from $2,000/mo. The audit exists partly so you can judge the work before paying for it — you keep the scorecard and the plan either way.

If it is useful, our comparison of ten agencies in this space — including ourselves, with the affiliation disclosed and the criterion where we rank last named — is in best GEO agencies in 2026.