Why Choosing a GEO Agency Is Hard
Choosing a GEO agency comes down to telling reproducible measurement apart from a good-looking screenshot: shortlist the ones that baseline on real devices first, use a method you can re-run, can show which citation sources actually appeared in tested answers, and whose contract contains only deliverables and measurement terms.
GEO is a young market. As a buyer, you may run into proposals that use similar vocabulary, and you may be shown a single screenshot as evidence: an AI answer that recommended a client once. The trouble is that AI answers drift by nature. The same question can return different answers across rounds, accounts, and phrasings. A screenshot proves something happened once. It doesn't prove it repeats, and it doesn't prove optimization caused it.
This guide breaks evaluation into seven dimensions, then compresses them into a 7-question checklist you can take into shortlist calls. If GEO itself is new to you, start with What Is GEO, then come back to the buying question.
Baselines and Methods: How They Measure Before They Optimize
The first thing to ask any GEO agency: show me how you measure, before we talk about how you optimize.
The difference between a baseline and a demo fits in one sentence: a demo can pick its moment and its phrasing, a baseline cannot. How question banks are designed, environments controlled, and rounds repeated is covered in full in How to Measure AI Visibility; we won't repeat it here. What procurement actually needs to verify is whether the method can be audited by you:
- Can raw records be spot-checked: per-question records (question text, answer, displayed citation sources, timestamp) kept on file, with sample pages available on request.
- Are the question bank and environment versioned: bank version, test environment, and the app or endpoint tested all logged, so results before and after stay comparable.
- How are failures and fluctuation reported: when a round shows nothing, or shows a regression, does it go into the report as observed, or do only the good rounds surface.
- Is the clean-environment boundary stated: clean environments reduce account-history variables; they don't mean every user sees the same answer. Writing that boundary into the report is itself a signal of measurement discipline.
For a sense of scale: for one global eKYC brand (a B2B SaaS), the monitoring question bank ran to 54 questions, frozen and re-tested on a schedule.
Citation Attribution: Do They Know Who AI Cites?
An agency that can't tell you which citation sources actually appeared in tested answers for your category can't tell you where your content should go.
The attribution method itself (aggregating the sources that recur across answers and prioritizing placement by that distribution) is covered in What Is GEO and How to Measure AI Visibility; we won't expand on it here. What procurement should verify is three things:
- Can they deliver an auditable source map: sources listed item by item, each traceable to a specific question and round, not a diagram nobody can check.
- Do they separate displayed citations from inferred sources: the two carry different evidential weight, and mixing them distorts attribution.
- Are versions and raw evidence retained: cited sources change over time, so conclusions should carry collection dates and versions, with original answers kept on file.
An agency that can hand over all three has earned the "where should content go" conversation; one that can't is not proposing from data.
Content Supply and Engine Coverage
GEO delivery capability shows in whether an agency can cover the kinds of places that actually appeared as citations in tested answers, supplying content written natively for each platform.
An agency that only writes for your own site can't cover the citation sources that show up in testing. Beyond your site, look for at least three kinds of ground: authoritative third-party records, high-intent Q&A, and vertical communities. Each has its own native language, and posting one identical draft everywhere means being slightly wrong on every platform.
Engine coverage follows the same logic, chosen by your actual target markets: for the Chinese domestic market, Doubao, Qwen, and DeepSeek can be tested; for overseas markets, ChatGPT, Perplexity, and Gemini. In our own testing we observed that the citation source types displayed in answers differ between the two sides, so measurement and content supply have to be handled separately. If your market spans both and a plan tests only one side, the other target market simply goes unmeasured. Ask which engines get measured, and how each side is tested and supplied.
Promise Boundaries and the Itemized Ledger
Be wary of guaranteed rankings and guaranteed sales. Nobody directly controls what a generative model says, so a contract should contain only a deliverable list and measurement terms.
This is the fastest single filter. Outcomes like a top position or a sales figure are not under any agency's unilateral control, so ask for the measurement terms and responsibility boundaries in writing. The honest posture: state what gets delivered, define how it gets measured, and report results as observed, gains and regressions alike.
Our own practice is one reference point: the contract sets out question-bank size, retest cadence, per-question screenshot logs, and a report you can reconcile. Changes in answer-list inclusion are reported as measured, with no promised lift, and no ranking or sales promises.
Past the promise question, look at the ledger. GEO produces a high volume of process output: what was published, where, when, and what each measurement round covered. The workable standard is an itemized ledger the client can check line by line at any time. Ask to see a sample page. For a fully documented engagement that went from absent to named, see this SaaS case study.
Red Flags Worth a Follow-Up Question
The patterns a buyer may run into all map back to verification items covered above: screenshot-only evidence points back to spot-checking per-question raw records; ranking or sales guarantees point back to measurement terms and responsibility boundaries; summary-only numbers point back to ledger audit rights; a one-off test with no mention of fluctuation points back to how failures and drift get reported. For content-side patterns such as one draft posted everywhere, see the distribution section of What Is GEO; we won't repeat it here. None of these calls for a verdict on the spot: bring out the matching verification item and ask.
The 7 Questions to Ask Every Shortlisted Agency
The seven dimensions, compressed into 7 questions you can ask in a call. Ask every candidate, write the answers down, and compare.
- 1. Before proposing anything, can you baseline my brand on real devices? Good answer: yes, with per-question records.
- 2. How is your question bank designed, and is the environment neutral and multi-round? Good answer: derived from the buyer decision chain, clean environments, multiple rounds, a frozen set for re-tests, failures and fluctuation reported as observed.
- 3. In tested answers for my category, which citation source types actually appeared? Good answer: an auditable source map, with placement priority following the observed distribution.
- 4. Beyond my own website, where can you supply content? Good answer: third-party records, high-intent Q&A, vertical communities, each written natively.
- 5. How do you cover domestic Chinese and overseas engines? Good answer: matched to your actual target markets, measured and supplied separately on each side.
- 6. What does the contract promise? Good answer: only a deliverable list plus measurement terms — question-bank size, retest cadence, how records are kept, how the report is computed; treat guaranteed rankings or guaranteed sales as a warning sign.
- 7. How do I audit the work? Good answer: an itemized ledger, open to line-by-line spot checks.
An agency that welcomes these questions is telling you something. So is one that steers around them.
FAQ
- How do I choose a GEO agency?
- Evaluate seven things: real-device baselines, a reproducible measurement method, citation-source attribution, content supply beyond your own site, honest promise boundaries, an itemized ledger, and coverage of the engines your market actually uses. The fastest test is asking for a baseline before a proposal.
- What should a GEO agency guarantee?
- Nobody directly controls generative AI output, so guaranteed rankings and guaranteed sales fall outside honest scope. A contract should contain only a deliverable list and measurement terms, with results reported as observed.
- Why isn't one AI screenshot proof of GEO results?
- AI answers fluctuate: the same question can return different answers across rounds, accounts, and phrasings. A single screenshot proves one occurrence; credible evidence is per-question records from multi-round testing against a frozen question set.
- Should a GEO agency cover both Chinese and overseas AI engines?
- Choose by your actual target markets: for the Chinese domestic market, Doubao, Qwen, and DeepSeek can be tested; for overseas markets, ChatGPT, Perplexity, and Gemini. In our testing, the citation source types displayed in answers differed between the two sides; a plan that tests only one side leaves the other target market unmeasured.
Shortlisting GEO agencies? Take these 7 questions to every one of them, including us. Or start with a real baseline test of your brand using public information, and decide from the data.
Get your free Growth Audit