Every AI visibility number we report is measured over a prompt set. The set is the list of questions we run through ChatGPT, Gemini, Google AI Overviews and the other engines every day, and the visibility figure is the share of those answers that name the client. Change the questions and the same brand reads 5% or 60% with nothing changing in the market. That is why we treat the prompt set as the most important thing we build in the first month, and why we are open about how it is built.
OpenAI, Google and Anthropic do not publish what people ask their assistants, and there is no keyword tool for ChatGPT. So an agency either guesses what buyers type, or it gets as close to the buyer as it can and reconstructs the questions from evidence. We do the second, on every client we onboard, in the order below.
Where the prompts come from
We pull buyer language from four places before a single prompt is written. Each one shows a different side of the same buyer, and a theme has to show up in more than one of them before it earns a place in the set.

Google data
The first source is the client's own Google footprint. We pull Search Console, GA4, Semrush and Ahrefs and read three things: which questions already bring people to the site, which questions in the category carry volume, and where the client is being found or missed. This is the only source with reliable numbers attached, so it anchors the rest.
Search Console has become more useful than it looks. When an AI engine answers a question, it runs its own web searches in the background, and many of those searches land in Google. They show up in Search Console as long, fully formed questions with impressions and almost no clicks, because the searcher was a machine building an answer.

On one client's property over four months, 273 of 498 queries were six words or longer. Those 273 queries carried 97,482 of 195,183 impressions and produced three clicks. The top one, "what is the best app for abandoned cart flow marketing automation?", had 21,151 impressions at position 5 and not one click. Nobody types that into Google and then ignores every result. An engine did, on a buyer's behalf, and Search Console kept the record. We read that layer as a log of how AI engines phrase the category, straight from the client's own data.
The client's own conversations
The second source is the client's inside knowledge, and it is the one most agencies skip because it takes work to collect. Buyers say what they want in support tickets and on sales calls, in the questions the same two or three people in the company answer every week. That language is closer to a real prompt than anything on the website, because the website is written in the company's words and the tickets are written in the customer's.
We collect as much of it as the client has. Support ticket exports and the frequently asked questions the team keeps. Recorded sales calls, when the client uses a notetaker. And an interview we schedule with the most customer facing person in the company, whoever hears the questions first: the head of support, or the salesperson who takes every inbound call. We record that interview and use it directly in the prompt work.
The client does not have to prepare anything. We connect to whatever tools they run, export, transcribe and sort. The one thing we ask for is the time of the person who talks to customers, and access to the data. Everything we take is stored under the client's data agreement and used only to build their set.
What comes out is a ranked list: every question a prospect or customer raised, how often it came up, and the exact words they used. On a large client this is hundreds of calls and thousands of questions. On a small one it is one good interview. Either way, it is the closest thing to a ChatGPT search log that exists for that business.
Public buyer conversations
The third source is where the client's buyers talk to each other. We start from the ideal customer profile and ask where those people ask questions in public: Reddit, Quora, Facebook groups, marketplaces, industry forums, LinkedIn threads, review sites. Then we go into those places and collect what they ask, how they phrase it, what tone they use, which topics keep coming back and which brands they mention to each other.
This source does two jobs. It fills the gaps the client cannot see, because a person who never became a lead never entered a ticket. And it sets the vocabulary. A buyer on a forum writes "my email flow isn't recovering carts, what else can I do", while the website says "omnichannel revenue automation". The prompt has to use the first one.
The engines' own sub-queries
The fourth source is the AI engines themselves. When ChatGPT, Gemini or Google AI Overviews answer a question, they issue a series of web searches to gather material, and our tracking captures those searches. Reading them tells us two things no other source can: how the engine rephrases a buyer's question before it searches, and which brands it adds on its own.

For one tracked prompt, "AI tools for abandoned cart recovery e-commerce", ChatGPT ran eight sub-queries in a week. Three of them named competitor brands the buyer never mentioned, and three went straight to specific competitor domains. Across the whole 100-prompt set that week the engines ran more than 1,200 sub-queries. The brands that keep appearing in them are the real competitive set for the tracking, whatever the client's own competitor list says, and the wording is a direct sample of how the engine understands the category.
The rule that joins the four
A theme that appears in only one source is a hypothesis. A theme that appears in three or more is a prompt cluster. We fuse the four sources into one table, one row per theme, with the count from each source beside it, and the clusters with the most independent support get the most prompts. A topic that only the founder cares about, and no ticket, forum or engine query supports, does not get tracked.
How a prompt is written
Once the demand table exists, the prompts are written from it using our own tooling and then edited one by one by a person. The generation is fast. The editing is where the set is won or lost, because a prompt that reads well to a marketer is often a prompt that no buyer would type.
The shape of a real question
Real questions are short and carry one piece of context. Look at any Search Console export, any keyword tool's question list, any sales-call transcript: people write "what's the best tool for X", "is there an AI that does Y", "does Z work with Shopify", "how much does this cost". They give one detail about their situation, the platform they use, the size of their store, the stack they already run, and then they ask.

So our prompts follow that shape. Most are 6 to 20 words. Each carries at most one context detail. Each opens the way people open: what's the best, is there a tool that, which one, how much, does it work with. And each ends in an ask that makes the engine name vendors, because a prompt that only describes a problem gets an answer full of advice and empty of brands, and there is nothing to measure in that.
A few things we cut on sight. Job-title preambles, because nobody introduces themselves to ChatGPT. Invented numbers stacked into one line, because a buyer gives one detail, never five. Deadlines pasted into the question, because a person choosing a tool is not ordering food. And prompts that doubt their own premise, "is it even worth adding calls", because the buyer already knows why they are asking.
Unbiased wording
The prompt is written in the buyer's words and never in the client's. This matters more than it sounds. If a company calls itself a "kitchen studio" and the prompts say "kitchen studio", the engines will find that company more often than they would for a buyer who types "kitchen furniture maker", and the visibility number will be measuring the wording. The market has not been asked. The same applies to coined product categories, internal feature names and any phrase that exists on the landing page and nowhere else.
We check every prompt against the vocabulary bank from the four sources. If the phrase came from a ticket, a forum thread, a Search Console query or an engine's sub-query, it stays. If it came from the brand book, it goes.
Coverage across the buyer's journey
A buyer does not ask one question. They ask a chain of them, from "why are my carts not recovering" through "what tools do this" to "does this one work with my stack and what does it cost". A set that only covers the last link measures the client at the moment of purchase and misses everything before it, which is usually where the client is least visible.

So the set is planned across the journey. Top-of-funnel prompts state the problem and ask what solves it. Mid-funnel prompts ask for the best options in the category, generically and for a specific type of buyer. Bottom-funnel prompts ask about fit, integration, price and comparison. On top of that we vary the persona where the client has more than one, and we build a separate set per market and per language, because the same question returns a different answer in Germany and in the US, and a different one again in German.

The mix is decided by where the client's revenue is, never by an even split. One recent set of 222 prompts runs 32 top-of-funnel, 113 mid-funnel, 57 bottom-funnel and 20 branded, across 12 topics, because for that business the mid-funnel is where the decision is made and the tracking should be densest there.
How many prompts
The total follows from the number of segments, so we count segments first.
We list the topics that matter, then the journey stages, then the personas and markets the client serves, and multiply. Each segment gets five to ten prompts, because fewer than five and a single strange answer moves the number, and more than ten and the segment is being measured twice. A single-market business with seven topics and three stages has 21 segments, which lands between 105 and 210 prompts. A business selling in nine countries lands in the hundreds. An e-commerce brand with many product categories and both awareness and purchase intent can run into the thousands, and we track sets of that size.

So the sets we run start at 50 prompts for a focused single-market business and grow to 100, 200, 500 and beyond as regions, personas, markets and product lines are added. The number comes out of the coverage plan.
Branded prompts
Every set we track includes branded prompts, and they live in their own lane. A branded prompt names the client or a competitor: "what do you know about X", "X vs Y for a Shopify store", "is X good for a small team". The client is named in the question, so the client is named in the answer nearly every time, and if those prompts were mixed into the visibility score, the score would rise without the market changing. So they are tagged separately and never counted in the headline number.
They are tracked because they carry data nothing else gives us. How the engines describe the client, in what tone, and to whom they recommend it. How the client comes out in head-to-head comparisons with each competitor. And whether something new has landed: when a client ships a feature, opens a store or changes a policy, we add branded prompts about it and watch how long the engines take to notice and whether they describe it correctly. Reputation in AI search is a workstream of its own, and the branded lane is where it is measured.
Topics and tags from day one
Every prompt is filed under a topic and tagged with its stage, persona, market and type before tracking starts. The structure has to be right at the beginning, because a tag added in month four cannot be applied to the history behind it. When a report later says the client is invisible in one topic on one engine, it is this structure that makes the sentence possible.
Testing before the set is confirmed
A prompt set is never confirmed on paper. Every candidate prompt passes two tests, and the set that goes to the client for approval is the one that survived both. This step is where most of the cuts happen, and it is the reason we can stand behind the number later.

The first test, by hand
We run every prompt two or three times in ChatGPT, Perplexity and Google AI Overviews and read the answers. Three questions. Does the engine give a relevant answer that names brands? Does it search the web to build it, with sources cited, or does it answer from memory? And could the client plausibly appear in that answer, given the sources the engine used?
Prompts fail this test in two ways. Some come back as advice with no vendor named, which means the engines read the question as a request for tactics; those are rewritten until the ask is unmistakable, or cut. Others come back from the model's training memory with no web search at all, which happens on well documented topics where the model already holds an answer. Those we cut, because no amount of work on the client's site changes an answer the model is not looking up. This check asks one thing: whether the prompt can be moved at all. Whether the client appears in it yet is a separate question, and the trial week answers that one.
The second test, seven days of tracking
The surviving prompts go into a trial tracking project for seven days, run daily on the client's engines, before anything is confirmed. Seven days gives seven answers per prompt per engine, about twenty across three engines, and that is enough to see the shape.

We read four things from that week. Visibility per prompt, for the client and for every competitor. The split by engine. The domains that carry the answers, which tells us where the citations come from. And the brands the engines add without being asked, which usually extends the competitor list.
Then both ends of the distribution get reviewed. At the bottom are prompts where no brand at all is named across a week of answers, and those are cut, because they cannot be won by anyone. At the top are prompts where the client already sits above 30%. Those get a hard look, because a set full of questions the client already wins reads well in the first report and cannot move afterwards. We keep them when the evidence shows buyers genuinely ask that way, and swap them when the wording was tilted toward the client's strengths.

The tilt check is systematic. We score every prompt on eight kinds of bias: whether the wording echoes the client's own marketing, whether it leans on the channel or feature where the client is strongest, whether it sits in the client's home language and geography only, whether it clusters at one funnel stage, whether it speaks to the client's favourite persona or vertical, whether it lacks the competitor framing real buyers use, and whether the wording is a year out of date. A set that scores high across those is measuring the client's positioning, and it gets rebalanced before the client sees it.
This is also the point where the trial data resolves the question the first test could not: which engines the client can move. On one 100-prompt set the same brand read 15.5% on Gemini, 13.5% on Google AI Overviews and 3.3% on ChatGPT in the same 30 days. The average, about 11%, would have hidden the engine where the work was not landing. Every read we do from this point on is split by engine for that reason.

The client review
After the two tests we present the set to the client: the prompts, grouped by topic and stage, with the evidence behind each cluster and what the trial week showed. The client reads it as the people who know their buyers best, and that read is the last input before tracking starts.
Clients usually change two things. They correct a phrase that a real customer would never use, which is exactly the check we want from them. And they flag a topic we weighted wrong, because they know a segment is growing that the historical data does not show yet. We adjust what can be adjusted, re-test anything that was rewritten, and confirm the final set in writing.
From that point the set is frozen. No prompts are added, removed or reworded during the engagement, because the set is the ruler the whole engagement is measured with, and a ruler that changes length cannot show progress. When a client wants to track a new market, a new product line or a new segment later, that becomes a new tracked scope with its own set and its own baseline, reported beside the original as its own line.
The baseline
Tracking runs daily from the day the set is confirmed, and the first two to four weeks are the baseline: the client's visibility, the average position in the answer when they are named, the sentiment, the share of answers that cite the client's domain, and the domains that are cited instead. Every figure is recorded per engine and per topic, and the baseline is written down with its dates before any work on the site begins.
Two rules keep the baseline honest. The prompts run from clean, logged out sessions, because an engine that has seen someone ask about a brand before will name that brand more readily, and a founder checking their own company from their own account sees a rosier picture than a stranger does. And the baseline is a measured number, never an assumed one. A client we were told sat at zero measured 3.2% once the first weeks of data came in.
The baseline is the reference for everything after it. Each month we compare against it, per prompt, per topic and per engine. Because the prompts run every day and every answer is stored, we also see the engines change underneath us: a forum that starts being cited across a category, a type of page that stops appearing, an engine that begins to search the web on a question it used to answer from memory. Those shifts are caught within days, and they feed directly into what we build next for the client.
The method in one list
Pull the client's Google data: Search Console, GA4, Semrush, Ahrefs. Isolate the six-word-plus queries.
Collect the client's own buyer conversations: support tickets, FAQ, recorded sales calls, and one recorded interview with the most customer facing person.
Map where the ideal customer talks in public and collect the questions, phrasing and tone from there.
Pull the engines' own sub-queries for the category and note the brands they add unprompted.
Fuse the four into one demand table. A theme needs three sources to become a prompt cluster.
Write the prompts from the table: 6 to 20 words, one context detail, buyer wording, an ask that makes the engine name vendors.
Plan coverage across journey stages, personas, markets and languages. Size each segment at five to ten prompts and let the total follow.
Write the branded and comparison prompts into their own lane, outside the visibility score.
Assign topics and tags to every prompt before anything runs.
Test every prompt by hand in ChatGPT, Perplexity and Google AI Overviews. Cut what returns no brands or no web search.
Run the survivors for seven days in a trial project. Cut prompts no brand wins. Review prompts the client already wins. Score the set for bias and rebalance.
Present to the client, take their corrections, re-test the changes, confirm in writing.
Freeze the set. Track daily. Record the baseline over two to four weeks, per engine and per topic, before any work starts.
Questions we get asked about this
How many prompts do you track for a client?
It depends on how many segments the business has. Sets start at 50 prompts for a focused single-market business and run to 100, 200, 500 and into the thousands for brands with many markets, personas or product lines. The count follows from the coverage plan.
Do you use AI to write the prompts?
Yes, with our own tooling, and every prompt is then edited by a person against the four sources. The generation is the fast part. The value is in the evidence the generator is fed and in the two tests every prompt has to pass afterwards.
Why are branded prompts kept out of the visibility score?
Because a prompt that names the brand gets an answer that names the brand almost every time, and mixing those into the score inflates it without anything changing in the market. We track them separately for what they do show: how the engines describe the client, how it fares in comparisons, and how fast a new feature or policy is picked up.
Can we add prompts after tracking starts?
The confirmed set stays frozen for the length of the engagement, because it is the ruler every result is measured with. New markets, products or segments are tracked as a new scope with its own set and baseline, reported alongside the original.
How long until we have a baseline?
Two to four weeks of daily tracking after the set is confirmed. The baseline is written down with its dates, per engine and per topic, before any work on the site begins.
What if we already rank well on most of the prompts?
Then the set is probably tilted toward your strengths, and we would rather find that out in the trial week than in month three. Prompts you already win are kept only when the evidence shows buyers really ask that way. The rest are swapped for the questions where the gap is.
Which tool do you track with?
Peec AI. Agenzy is an official Peec AI partner. The prompt method above is ours and works with any tracker that runs prompts daily and stores every answer.




