How to Evaluate an AI Search Optimization Agency Without Buying Hype
A buyer-side framework for AEO, GEO and AI-search services: what is genuinely new, what overlaps with SEO, what evidence to request and which promises to reject.
Answer engine optimization, generative engine optimization and AI search optimization are being sold as new categories. Some parts are genuinely new: prompt-set monitoring, citation analysis and measuring visibility across systems that do not behave like ten blue links. Much of the underlying work remains familiar: crawlability, clear information architecture, factual content, entity consistency and earning references from other sites.
Google's current guidance says the same foundational SEO practices apply to its AI features and that there are no additional technical requirements or special schema needed to appear. That makes procurement simpler: an agency should be able to explain the new measurement layer without pretending the web was reinvented. See Google's AI features guidance (opens in a new tab) and generative AI optimization guide (opens in a new tab).
What an AI-search engagement should actually contain
| Workstream | Useful output | Weak substitute |
|---|---|---|
| Baseline | Versioned prompt set, engines, markets and current citations | One screenshot of a favourable answer |
| Technical access | Crawl, render, index and canonical validation | A proprietary 'AI readiness score' |
| Content evidence | Claims mapped to sources, expert review and update ownership | Rewriting every paragraph into question-and-answer format |
| Entity consistency | Reconciled organisation, product and author facts across owned profiles | Adding unsupported schema fields |
| Authority | Plan to earn relevant third-party citations | Bulk placements described as 'LLM seeding' |
| Measurement | Prompt-level visibility, citations and assisted commercial outcomes | A single visibility percentage with no reproducible inputs |
The prompt set is the measurement contract. It should state the questions, audience, market, language, systems checked and collection date. Without that versioned input, a visibility score cannot be reproduced and trend lines can be improved by quietly changing the questions.
Ask what changed from the agency's pre-AI process
A credible provider can show which workflows changed and which did not. Seer Interactive's public SEO RFP guide recommends asking agencies to compare a pre-AI and post-AI process and explain how they validate changing search behaviour. That is a better test than asking whether the agency 'does GEO.' Review Seer's RFP questions (opens in a new tab).
Good answers usually add prompt research, citation tracking, entity reconciliation and new reporting. Weak answers rename content briefs, technical audits and digital PR while adding an unexplained platform fee.
Evidence to request
- 1A redacted baseline showing the exact prompt set and collection method.
- 2A citation change linked to specific work, with collection dates and screenshots.
- 3A case where visibility did not improve and the hypothesis changed.
- 4The split between branded prompts and non-brand buying questions.
- 5How the team handles answers that vary by geography, account state or repeated runs.
- 6Which commercial outcome is measured after visibility: assisted visits, branded demand, qualified leads or revenue.
Do not accept a case study that reports only the number of prompts where a brand appeared. Branded prompts are easy to win and often measure awareness created elsewhere.
Red flags
- Guaranteed citations in ChatGPT, Gemini, Perplexity or Google AI Overviews.
- Claims that a secret schema type or hidden file is required for inclusion.
- Hundreds of machine-written pages created only to restate the same facts.
- No separation between classic organic visibility and AI-answer visibility.
- A proprietary score whose prompt set, sampling schedule and engines are hidden.
- Recommendations to publish facts without a named source or subject-matter review.
How to build a defensible shortlist
The directory does not label a company an AI-search specialist unless its profile documents that capability. Until that evidence exists at useful scale, start with the adjacent capabilities the work requires: technical SEO, content marketing, digital PR and analytics. Then use the evidence questions above to verify the new layer directly.
AI-search procurement should increase the standard of proof, not lower it. The systems are variable, measurement is immature and guarantees are especially easy to manufacture. Buy a reproducible process and documented judgement, not a new acronym.
SEO Companies Hub Editorial
Directory Research
SEO Companies Hub publishes independent, research-backed guidance for US businesses choosing an SEO partner. We score agencies on published facts, separate marketing claims from evidence, and date-check anything that can change.