Choose a GEO agency on two asks that a proposal cannot fake: a dated case study with measurable AI citations at both ends, and one monthly report that splits results per engine and connects them to leads. Ask for last month's report with the tracked prompt list beside it before the next call, and treat that single request as the first filter. Eleven questions then carry the call itself.
Across 22 agencies whose public pages, directory profiles and press coverage we reviewed in September 2026, 17 of 22 sold an AI visibility audit, 15 sold citation-ready content, 11 sold technical and schema work, and 10 sold digital PR. The stack is the same, the words are the same, and the price is usually missing.
What decides whether you get anything sits underneath the deliverables. It starts with who writes the tracked prompt list and when you approve it, whether the starting number was measured or assumed, and whether anything is held back untouched so a rise can be traced to something you paid for. After that come who does the off-site work, who implements the technical changes, what happens at month three and what you keep if you leave.
The 22-agency review is global and counts only what each agency publishes about itself. The best GEO agencies in the US in 2026 is narrower: eleven agencies a US buyer meets, Agenzy included, scored one by one on the same eleven criteria, with a source link and a check date in every cell.
Why every GEO proposal reads the same
GEO proposals read the same because the deliverable stack is identical across the market, and the business model underneath it never appears in the deck. Of the 22 GEO and AEO agencies we reviewed in September 2026, eight are North American, four British, seven continental European and three Asian.
The table below counts what each of the 22 agencies names as a deliverable on its own public pages, read in September 2026.
What the proposal offers | Agencies naming it on their public pages (of 22 reviewed, September 2026) |
|---|---|
AI visibility audit | 17 of 22 name an AI visibility audit |
Citation-ready content | 15 of 22 name citation-ready content |
Technical and schema work | 11 of 22 name technical and schema work |
Digital PR and off-site citations | 10 of 22 name digital PR or off-site citation building |
A tracking platform of their own | 14 of 22 name a tracking platform of their own |
Entity and knowledge-panel work | 4 of 22 name entity and knowledge-panel work |
Reddit or other community participation | 2 of 22 name Reddit or other community participation |
Training for the client's team | 2 of 22 name training for the client's team |
A monthly price published on their own site | 4 of 22 publish a monthly price on their own site |
Almost everyone sells the same four things, most claim a platform of their own, and four publish a price. Two proposals are therefore indistinguishable until you have sat through two calls.
The platform claim needs a second look. Fourteen of the 22 named a tracking product of their own, and at least two of those are an interface built over a licensed tracker. Licensing is a reasonable way to build a product. What decides your reporting is which engines the underlying data covers and how often the prompts run.
Behind the identical stack sit four different businesses. Three things in the material you already hold will sort a deck into one of them. The first is how much of the agency's own navigation belongs to GEO. The second is whether the headline case metric is traffic and rankings or AI citations, and the third is whether the tracking product is described as a license or as their own build.
Each tell points at a row in the table below. Navigation share separates full-service from specialist: GEO as one line among ten services is the first row, GEO owning the menu is one of the other three.
The headline case metric then splits the specialists, since traffic and ranking numbers mark the SEO extension sitting inside a full-service retainer and AI citation numbers mark the content-led shop. A tracker described as a licensed feed usually puts the agency in the technical specialist or the entity and reputation row, since both buy their measurement in and spend the month on something else.
The table below sorts the same 22 public pages into four business models and says what each one changes on your account.
Business model | Agencies in this model (of 22 reviewed, September 2026) | What it changes on your account |
|---|---|---|
Full-service or enterprise, GEO inside a wider media or SEO retainer | 11 of 22, the most common model in the review | A slice of a large team's month, and a budget that can move back to media |
Content-led specialist | 7 of 22 | Writers and editors at the centre of the work |
Technical specialist | 2 of 22 | Crawling, structure and schema, with content left to you |
Entity and reputation | 2 of 22 | Brand properties and how models describe you, ahead of publishing volume |
The model decides who sits on your account, and the deliverable list never tells you which one you are buying.
The buyer guides repeat the same shape. On 2026-09-16 we read four agency-published pages that answer this question, published between 2026-05-06 and 2026-08-04. Two of their points hold up: demand pipeline outcomes over visibility scores, and ask which AI surfaces the agency monitors and why. None of the four mentions a baseline or a holdout, and none asks how results are reported per engine, who owns the prompt list, or what the term and exit look like.
Is GEO worth paying an agency for, or is it hype?
GEO is worth paying an agency for when three things are true at once, and it is hype whenever the retainer has no way to show its own effect. The three: buyers in your category ask assistants for recommendations before they shortlist, your site can be crawled and changed, and somebody on your side can attach the result to leads.
Where those hold, the fee buys two things you cannot assemble piecemeal. The first is measurement: a fixed prompt set read daily in a tool with a baseline behind it. The second is the off-site half of the work: the third-party pages and community answers the engines already trust, which take longer to earn than anything you publish yourself.
The hype version is easy to name. It reports a visibility percentage with no baseline behind it, no holdout beside it and no lead line under it, so the number can rise for reasons nobody can separate from the work.
Hire a GEO agency when the off-site half and the pace are the bottleneck, and build in house when the content and technical layers are. Building in house needs a developer who can change robots.txt, schema and page structure, somebody publishing every week, and a tracking subscription. Most marketing teams can staff that.
Placements, community answers and the third-party pages engines already trust run on relationships and a weekly rhythm, and that is the half teams keep postponing until a quarter has gone by with nothing measured. The hybrid arrangement is common: the agency runs measurement and the off-site work while your own team writes.
There is also a case where the right answer is no. If nobody in your category asks an assistant for a recommendation, no agency can manufacture that demand. Where the demand exists, buy the measurement first, since measurement is what tells you whether the rest of it worked.
In the engagements we have measured, month one goes on instrumentation and answers usually begin to shift in month two or month three, soonest where the starting point was lowest. A proposal worth signing says as much on paper: month one instrumented, movement given as a range of months, and no promised position or percentage anywhere in the document. The full sequence sits in GEO for B2B SaaS in the first 90 days.
Two asks that are hard to fake: measured client results with both ends of the number
Two asks separate an agency that does this work from an agency that describes it well: one dated case with AI citation numbers at both ends, and one real monthly report with the engines split out and leads attached.
A case study passes when it names or describes the client, gives the number before and the number after, states the window, and names the instrument that measured it. Jessica Ehrhardt put the first half of that test into a buyer question published on 2026-08-25: can they show "a case study with measurable AI citations or AI-platform referral traffic".
A percentage alone fails it. "340% more citations" is unfalsifiable without the starting count, and a screenshot of one ChatGPT answer proves only that an answer existed on one day. Three of the 22 agencies in our review published a named client with a before-and-after AI metric attached. That is a count of published material, since a public page is not a delivery record, and the same limit applies to every count from the review.
A monthly report passes when you can see four layers: answer coverage per engine, the URLs the engines actually cited, what the site did with that traffic, and leads by source. Ask for one from any client, with the numbers masked if they have to be. A report like that takes weeks of daily prompt runs to exist at all, which is why it cannot be assembled during a sales cycle.
We publish an annotated example in what a GEO agency's monthly report should show, and a full diagnostic audit of one B2B SaaS company in what an AI visibility audit of a B2B SaaS company finds.
Baseline and holdout, the two mechanisms behind a GEO agency's numbers
A baseline is the measurement of your AI visibility taken before any work starts, on the same prompts, in the same tool, over at least two weeks. It is the only number later results can be read against. Ask when it will be measured and what it covers.
A holdout is a set of tracked prompts or pages left deliberately untouched while the rest of the program runs, so the two can be compared at the end. It answers the question every retainer eventually faces: would this have happened anyway. An engagement without one can show movement and can never show cause.
Both are cheap to ask for and awkward to answer. The version that survives scrutiny registers the test before the work ships, with the success threshold written down in advance, because a threshold chosen afterward is a story. The catch is that a holdout withholds something from someone, so on a small prompt set the workable form staggers the work across topics instead.
Per-engine reporting is the third piece of the same instrument. A blended score averages engines that draw on different sources. ChatGPT, Google AI Overviews and Perplexity share only 2 of their top 10 cited sources when answering the same questions about the same brand, according to Agenzy's 2026 source study (n=109 business projects). A brand can hold its place in one engine and lose it in another while the average barely moves.
Eleven questions to ask a GEO agency before you sign
Eleven questions cover what a GEO proposal leaves out. Each one carries an answer that passes, an answer that reveals a renamed SEO retainer, and a follow-up that breaks a rehearsed reply. The eleven, in the order they work on a call:
How much of your revenue comes from GEO work?
Who writes the prompt list, and when do we approve it?
Where does an AI-sourced lead appear in our CRM?
What was the starting number in your best case study?
What can we read about your method before we sign?
Whose data sits behind the dashboard we will be reading?
What will you hand our team that we can use without you?
What have you tested that did not work?
What does your own tracked prompt set look like?
What does it cost, and is that price on your website?
What changed in how you work after the last model release?
They come from our own scoring rubric, from three buyer questions Jessica Ehrhardt published on 2026-08-25, and from Lars Lofgren's published bands on what agency money buys.
The questions have one limit. An agency running for two years has more published material than a good team that started this year, so weigh what you hear on the call above what you can find online.
How much of your revenue comes from GEO work?
The share of revenue that comes from GEO work is the first thing to pin down: whether GEO or AEO is the core offer, how many clients sit on that service line, and what adjacent work the agency turns down.
An agency with ten services on the menu, where AI search is one line item, is the full-service model, the most common one in the review. That model decides who is on your account and whether the AI search work survives the next quarter when the media budget moves.
Ask about the other side on the same call. A specialist shop is usually small, so ask how many people will touch your account and who writes when the writer is away. A ten-service agency has a bench and thin depth; a specialist has depth and a thin bench, and both are real costs.
A service menu with no client count behind it has already answered the question. Then ask how many clients are on that service line today, and which piece of adjacent work they turned down this year.
Who writes the prompt list, and when do we approve it?
A working answer on the prompt list gives a count, a tool, the date you sign off on the list, and the rule for what happens to that list afterwards.
The good version is built from your Search Console data, an hour with your salespeople, and the forums where your category argues. The set then freezes, because a list that keeps changing produces a trend line about the list itself.
A weak answer offers keywords, or offers to monitor AI mentions of your brand, which tracks questions that already contain your name and scores high by construction. The follow-up that settles it: ask how many proposed prompts contain your company name, and ask them to remove those.
Where does an AI-sourced lead appear in our CRM?
An AI-sourced lead should arrive in a named field in your own CRM, so ask which field carries it and who builds that field in month one.
The real answer names mechanisms: a "how did you find us" field on the form or the call script, AI referral sessions in analytics, landing pages tagged by source. It states the limit too, since some assistants strip the referrer and part of that traffic arrives as direct. Ehrhardt puts it as a buyer question: "How do they connect that visibility to pipeline, not just citation counts?"
A weak answer returns citation counts and a visibility percentage with nothing underneath them. The follow-up: ask who on their side does the CRM setup, in which week, and what they need from your RevOps person.
What was the starting number in your best case study?
The starting number in a case study is the figure the client began with, and without it nobody can check the result.
An answer that passes is specific in a dull way: a named or described client, both figures, the weeks between them, and the tool that read both ends. Watch for a percentage with no denominator, an unnamed client in an unnamed category, or a screenshot.
Then ask across how many prompts the figure was measured, in which weeks, and what the same prompt set does today. Our own worked example, with the method and both ends of the number, sits in how the Outcraft case was measured.
What can we read about your method before we sign?
A GEO agency should be able to send you a URL that explains how it builds a prompt set, how it earns citations off-site, and what it does in month one.
Published method is the only work sample you get before signing, so read it for procedure: the order of the technical pass, how content is chosen against prompts, what gets measured when.
The weak version describes outcomes ("we make your content citable") with no step in it, or points at a news blog about the latest model release. Ask them to open the page that says how a prompt set is built, and check whether it names any source of prompts beyond a keyword tool.
Whose data sits behind the dashboard we will be reading?
Ask which tracker your reports come from, how often the prompts run, which engines it reads, and whether you get a login.
Fourteen of the 22 agencies in our review name a platform of their own, and at least two of those sit on a licensed tracker. What matters is the source underneath, the run frequency and the engine coverage.
A poor answer describes proprietary data and then cannot say which engines it reads or how often. Ask what happens to the measurement data if you leave, and whether you can export it.
What will you hand our team that we can use without you?
The answer you want names something your own people can pick up and run: a documented way of building a prompt set, a working session with your marketing and sales teams, a report template, a study whose method you can repeat on your own brand.
An agency that teaches has written its process down, which means the process survives a staff change on their side.
A link to their blog is the thinnest version of this answer. A gated PDF is weaker evidence than one thing a client is still using a year later. Ask for one thing they handed a client that the client still uses, and who at your company it would be addressed to.
What have you tested that did not work?
Ask for one named test with a method and a null result, and ask how they decide that something worked at all.
Listen for a baseline, a window and a comparison. Every agency's case studies work, and only a shop with a measurement habit remembers the tests that went nowhere, which are the ones that stop you paying twice for a tactic.
An agency that treats this as a trick question and returns another success story has told you what its measurement habit is worth. Ask what they stopped doing this year, and what the evidence was.
What does your own tracked prompt set look like?
An agency's own tracked prompt set should come back as a number, a tool name, and a note on how many of those prompts contain the agency's own brand name.
Where the answer is none, the agency is selling an instrument it has never turned on itself. Where the answer is a set full of its own brand name, the measurement flatters itself by construction.
The follow-up: ask how many branded prompts sit in that set, and whether they are excluded from the headline number.
What does it cost, and is that price on your website?
Ask for the monthly figure, the term, and whether that number appears anywhere public.
Four of the 22 agencies we reviewed publish a monthly price on their own site. The outside view comes from the bands Lars Lofgren published on LinkedIn in 2026: at $5,000 a month you can start in an organic channel with limits on scope, $10,000 buys real progress in one channel, and $20,000 to $40,000 is where full-service quality sits. His advice is to avoid agencies under $5,000 a month.
A custom quote with no band attached and no term gives you nothing to compare. Ask what the same money buys at half the scope, since a real answer describes what gets dropped first.
What changed in how you work after the last model release?
Ask for one specific change, with a date, caused by a named model or product release.
The answer should be concrete: a retrieval behavior that changed, a page structure adjusted, a channel that started or stopped paying.
Generic enthusiasm about the pace of AI, with no date anywhere in it, is the tell. So is a method page whose last update predates the model releases it should be reacting to. Ask for the date of the change and what they were doing the week before it.
How Agenzy answers these eleven questions about choosing a GEO agency
The table below gives Agenzy's own answer to each of the eleven questions, checked on 2026-09-28 against our service pages, our contract terms and our own tracking account.
Question a buyer should ask | Agenzy's answer, checked 2026-09-28 |
|---|---|
How much of your revenue comes from GEO work? | GEO and AEO are the core offer. ChatGPT Ads management runs as a separate service, and SEO is sold only as a block inside a GEO plan |
Who writes the prompt list, and when do we approve it? | 50 buyer prompts per client in Peec AI, built from the client's search data, a conversation with their salesperson and the forums their buyers use; approved by the client in month one, frozen after approval |
Where does an AI-sourced lead appear in our CRM? | Set up in month one: AI-channel sessions, visibility on the tracked set, and leads by source in the client's own CRM, from a "how did you find us" field on the form, the call script or the checkout |
What was the starting number in your best case study? | Outcraft AI, 0.25% to 11.9% AI visibility across 100 tracked prompts over one quarter, measured in Peec AI, client named and published |
What can we read about your method before we sign? | On the blog and the service page: how a prompt set is built, what month one covers, what the reporting shows |
Whose data sits behind the dashboard we will be reading? | Peec AI, where we are an official Peec AI partner, with our own tooling around it: an audit engine, bot-access checks, prompt-set bias scoring, competitor content watchers |
What will you hand our team that we can use without you? | Two published studies with their method attached, and the 50-prompt list the client approves in month one, written from a conversation with their sales team |
What have you tested that did not work? | Each test is registered before the work ships, with the success threshold set in advance; the two studies carry their own method |
What does your own tracked prompt set look like? | 150 prompts tracked on Agenzy in Peec AI in September 2026, the same tool clients get, 12 of them branded and excluded from every headline number |
What does it cost, and is that price on your website? | From 5,000 EUR a month, ex VAT; the monthly price is published on the service page |
What changed in how you work after the last model release? | Every OpenAI, Google, Anthropic and Perplexity release is read the week it lands and the method updated |
Read down the right column as the shape of an answer that passes, since every line in it points at a page, a tool or a number you can open.
The studies are which sources AI engines cite, built on 2,283,923 citations, and the Lithuanian AI visibility study, covering 15 market categories in one country and published in Lithuanian. The case sits at the Outcraft case study and the price on the GEO service page.
Two more answers sit outside the eleven. The month-one technical pass lands as changes on the client's site, deployed by us. The term is six months, and we hold a results review against the baseline in month three.
How to tell a GEO agency from an SEO retainer with a new name
A relabeled SEO retainer cannot produce last month's report with the tracked prompt list beside it. That pair requires a prompt set running daily for weeks, an answer-coverage number per engine, and the cited URLs.
What comes back instead is a keyword ranking table, an organic traffic chart, a slide about schema, and a screenshot of an AI answer that mentions the client. Each of those is real work, and none of it measures whether a buyer asking an assistant about the category hears the client's name.
The red flag has three parts and it outranks any score you give an agency: no tracked prompt set, no attribution into the CRM, and no per-engine reporting.
Term, exit, and what happens at month three
A GEO contract runs six or twelve months in most of this market, and the length matters less than what the agreement says you keep when it ends.
In our 22-agency review, 12 months is standard at the premium end, six months is standard in the middle, and two of the 22 sell a diagnostic engagement first with no lock-in. The exit terms are where they differ.
Six things to settle in writing before you sign:
The notice period, and what happens to work in progress once you serve it.
Ownership of the prompt set, the reports and the raw measurement data, which should be yours.
Ownership of the content and schema deployed on your site, which should also be yours.
Who does the off-site work: their own team, a freelance network or a subcontractor. Ask who signs outreach carrying your brand name.
Who implements the technical pass: their developers, your developers, or a recommendations document you have to resource yourself, which is a different price for the same deliverable.
A defined review point where either side can still change course.
Points four and five are the ones a proposal almost never answers.
Six months is money committed in advance, so the review point is worth testing in the sales call. Ask each agency what their most recent client review decided, and what changed after it.
FAQ
What should I ask a GEO agency before I sign?
Ask for the prompt list you will approve, the baseline measurement date, one dated case with AI citation numbers at both ends, and a real monthly report with the engines split out. Then ask who owns the data and the content if you leave, and what the review point in the contract actually decides.
How do I know if a GEO agency is just doing SEO with a new name?
Request last month's client report together with the tracked prompt list. A renamed SEO retainer returns keyword positions, traffic charts and a screenshot of an AI answer. A GEO program returns answer coverage per engine on a fixed prompt set, the URLs the engines cited, and leads by source. If none of those three exist, the work is still SEO under a new label.
How long should a GEO contract be?
Six months is the working minimum and twelve is common at the premium end. Two of the 22 agencies we reviewed open with a diagnostic engagement and no lock-in, which is a reasonable way to start. Length matters less than three clauses: the notice period, who owns the prompt set with the reports and the raw data, and who owns the content and schema that end up on your site. The last two answers should name your company.
Who should own the tracked prompt set, us or the agency?
You should own it, along with the reports and the raw measurement data, and that belongs in the contract. Approve the list yourself before tracking starts, since those prompts define every number that follows. Once approved, leave it frozen. Prompts added mid-engagement move the percentage without changing what any engine says about you.
Should we hire a GEO agency or build this in house?
Build in house when the technical and content layers are your bottleneck and you already hold the people; hire an agency when the off-site half and the pace are the bottleneck. In house works with a developer who can change robots.txt, schema and page structure, someone publishing every week, and a tracking subscription. Agencies earn the fee on placements, community participation and the third-party pages engines already trust. A hybrid is common: the agency runs measurement and off-site, your team writes.
How long before GEO shows up in AI answers?
Month one goes on instrumentation: the prompt set, the baseline and the attribution wiring. Answers usually begin to shift somewhere in month two or month three, and they shift soonest for brands that start at the bottom of the scale. Treat any agency that promises a position or a percentage as a warning, since no agency controls how a model composes an answer.
About Agenzy
Agenzy is a GEO and AEO agency: we get brands named and recommended inside ChatGPT, Gemini, Google AI Overviews, Perplexity, Claude and Copilot. Agenzy is one of the agencies a buyer would score with these eleven questions, and our own answers are published with them. Behind those answers sit 500,000+ AI chats analysed, 15,000+ prompts tracked, 150+ audits completed and 1,000,000+ EUR generated for clients by AI search. GEO plans start from 5,000 EUR a month, ex VAT, and client cases with dates sit at agenzy.lt/case-studies.




