Prompt tracking, also called AI visibility tracking, means asking AI engines the same buyer questions every day. A tool runs those questions through ChatGPT, Gemini and Google AI Overviews, and records which brands each answer names. Any business whose buyers ask an assistant for advice can use it.
What prompt tracking means in plain words
Prompt tracking is mystery shopping, done every day by a robot. A mystery shopper walks into a shop, asks the assistant what to buy, and writes down which brands the assistant named. Prompt tracking does that inside AI assistants instead of shops.
The robot asks the same questions every morning. It reads each answer and records the brands in it. After four weeks you have a scoreboard showing how often the engines name you, and how often they name the company down the road.
Where prompt tracking sits in GEO
Prompt tracking is the instrument for two of the four steps we work through with every client in generative engine optimization (GEO).
Open the door. The AI companies' bots can reach and read your website.
Give them something worth citing, on your site and on the sites the engines trust. Your pages carry facts, numbers and sources an engine can lift.
Watch what gets cited. Which of your pages the engines pull in, for which questions, how often.
Count the recommendations. How often the engines name your brand, for which buyer questions, in what tone, and what that turns into.
Agenzy calls all four steps GEO. Prompt tracking is the measuring half of that list. Steps one and two move the number, and steps three and four read it.
How prompt tracking works
Whoever writes the questions decides the score, on ChatGPT and on every other engine, because the number is only ever an average over the questions you chose.

Write the questions the way buyers ask them. Short, one detail about their situation, and an ask that makes the engine name suppliers.
Run them every day on every engine you care about. The sessions run logged out, because an account that has asked about your brand before will name it more readily.
Record what came back: was the brand named, how high up, which websites the engine cited, and in what tone.
Record the same for your competitors. Their score is what tells you whether yours is any good.
Read the trend over four weeks, engine by engine.
Once it is set up, a tracking tool runs the questions, records the answers and builds the trend on its own. Writing the questions is the part a person has to get right.
One tracked question, drawn out
A single tracked question shows what a tool records, and the example below is an illustration with a made-up brand and a plausible answer.
A buyer types: "best payroll software for a 50 person company in Germany". ChatGPT searches the web, reads what it finds, and answers with four tools. Your brand is named third. The answer links a comparison article on a review site, a Reddit thread, and one competitor's own pricing page.

The tool writes down what it saw: named yes, position three, tone neutral, sources cited, competitors present. Do that every day for a month and you can see whether third is your normal place or a good day.
How prompt tracking differs from Google rank tracking
Prompt tracking records whether an AI answer names your brand, while Google rank tracking records the position of your page in a list of links. The table below sets the two side by side on what a marketing team usually wants to know.
What you want to know | Google rank tracking | Prompt tracking |
|---|---|---|
What is asked | A keyword, often two or three words | A full buyer question, usually 6 to 20 words |
What is measured | The position of your URL in a list of ten blue links | Whether an AI answer names your brand at all, and where |
What the result depends on | One engine, Google | Five or six engines that read different sources and disagree with each other |
How stable the reading is | The same keyword gives a near identical result an hour later | The same question can give a different answer an hour later |
What a single reading is worth | Usable on its own | Noise on its own, so the unit is a four-week trend |
A brand can hold position one in Google and appear in no AI answers in its category.
What AI visibility tracking measures
AI visibility tracking reports the same handful of numbers whatever tool you use. We track clients in Peec AI, and these are the definitions the tool uses.
Visibility is the share of answers that name your brand. Run 50 questions daily for a week on one engine and visibility is the share of those 350 answers that mention you. Every engine is counted separately. Some vendors and some of our own pages call it answer share.
Share of voice is your slice of all the brand mentions in your category. Visibility says whether you are in the room, and share of voice says how loud you are against everyone else in it.
Citations count how often an engine used one of your pages as a source. A brand can be named with no citation, and a page can be cited with no mention, so both are tracked.
Position is where your brand sits in the answer. First named and fifth named are different outcomes for a buyer reading on a phone.
Sentiment is the tone of the sentence you appear in. An engine can name you as the cheap option, the enterprise option or the one people complain about.
Most brands start lower than they expect. Agenzy's 2026 study of 187,810 AI answers found the median European SMB brand appears in just 5% of AI answers about its own product category (across 90 tracked market projects, from April to August 2026; a second cut of the same study counts 91 projects).
Why the same question gives a different answer every time
AI engines are not repeatable machines, so two identical questions asked an hour apart can return different brands. The engine runs fresh web searches each time and writes the answer word by word, and a small change early in that process changes the rest.
That wobble is why the questions go out every day. A week of daily answers gives a share you can hold against last week's share. A single screenshot cannot.
Engines also disagree with each other, and the gap is wide. Agenzy's 2026 study measured a median 5.1x visibility gap between a brand's best and worst AI engine, and 30% of visible brands were completely invisible on at least one engine (across 86 brands visible somewhere, from April to August 2026).
So every reading we report is split by engine, because a blended score can hide a strong engine sitting next to a dead one. Read every number one engine at a time.
What a tracked prompt set looks like after four weeks
A real tracked set starts at 50 buyer questions, and the five below are an illustration built to show the shape of a tracked prompt set, with invented readings for a made-up payroll software company. Each cell is the share of that week's answers naming the brand, and the last column splits week four between two engines. Each week is seven daily runs, so every share is a multiple of one seventh: 14% is one answer in seven, 43% is three.
Buyer question tracked daily | Week 1 share | Week 2 share | Week 3 share | Week 4 share | Week 4 by engine |
|---|---|---|---|---|---|
best payroll software for a 50 person company in Germany | 0% | 0% | 14% | 29% | ChatGPT 43%, Gemini 14% |
payroll tools that handle German and Polish payslips | 14% | 29% | 29% | 43% | ChatGPT 57%, Gemini 29% |
how much does payroll software cost for a small company | 0% | 0% | 0% | 0% | ChatGPT 0%, Gemini 0% |
payroll software that syncs with our accounting system | 43% | 29% | 43% | 57% | ChatGPT 71%, Gemini 43% |
is it cheaper to outsource payroll or buy software | 0% | 14% | 0% | 14% | ChatGPT 29%, Gemini 0% |

The accounting-system question is the strongest position and the one worth defending. The cost question never names any brand at all, so no supplier can win it and it may not belong in the set. Gemini trails ChatGPT on every row where the brand appears at all, which makes Gemini the job for the coming month.
What counts as a prompt, and what does not
A prompt is one full question in the buyer's own words, and it never contains your brand name. Buyers write "what's the best tool for X" or "does Y work with Shopify". They give one detail about their situation and then ask.
A question that names your company is a vanity question. The brand is in the question, so the brand comes back in the answer almost every time, and the score rises without anything changing in the market. We track those separately and keep them out of the headline number.
A question also has to be one the engine answers by searching the web, because an answer written from the model's memory moves for nobody. And it has to ask who supplies the thing, because a question about the law or the price gets answered with advice and names no companies at all.
The full method, including the sources we pull buyer language from and the tests every question passes, sits in how we build the prompt set we track for a client.
When prompt tracking is not worth doing
Prompt tracking is not worth paying for in three situations.
Some categories cannot carry fifty buyer questions worth asking. Fifty is the floor. Below it the set is too small to average out the daily wobble, and every reading jumps around for no reason.
A site with nothing fixed yet gives the tool nothing to measure. If the bots cannot reach your pages and there is no content worth citing, tracking will faithfully report zero for months. Opening the door and publishing something worth citing come first, and the tracking then measures whether they worked.
A team with no time to act on the data buys a dashboard nobody opens. Someone has to read the per-engine trend each month, pick the question that slipped, brief the page behind it and see whether the next month's reading moved. If no one owns that hour, you are paying a subscription and getting a report.
Which tools do this
Several tools run prompt tracking, among them Peec AI, Profound, Scrunch AI, Otterly.AI and Semrush. We use Peec AI on client work and Agenzy is an official Peec AI partner. We compare what they actually read and how often in the best AI visibility tools in 2026.
Where to start
Start with the questions, because the tool is the easy part. Ten questions run by hand tell you whether you have a problem. A tracked set that produces a trend starts at fifty, and the sets we run inside the GEO service grow from there with topics, buyer stages, personas and markets.
Write ten buyer questions in an hour with your salesperson. Use the words they hear on calls, and let none of the questions name your company.
Run them daily on the engines your buyers use and leave them alone for four weeks. That four-week reading is your baseline, and every later number is compared to it.
Pick the one question worth winning and fix the pages behind it. Then watch whether the number on that question moves.
The format of the page you fix in step three decides how much of that work comes back as citations. In Agenzy's 2026 study, comparison-style pages earned a 1.6x higher median citation rate from AI engines than homepages, yet homepages still absorbed 29% of all citations (across 2,730 top-cited URLs in 91 projects).
Reddit appeared among the top-cited sources for 84% of the 91 business projects in Agenzy's 2026 AI search study, more than any other domain on the internet. A comparison page of your own and an honest answer in the thread your buyers already read are the two cheapest things to try on that one question.
If you want the measurement and the work in one place, that is what we do on the GEO service, and the method behind each step is published in our guides.
Book a free 30 min call and we'll show you where you appear in AI answers today.
Questions people ask about prompt tracking
How many prompts should I track?
Fifty questions is the floor for a focused single-market business with one product line. Agenzy sizes a set by counting topics, buyer stages, personas and markets, then allowing five to ten questions for each one. Fewer than five questions per segment and one odd answer moves the whole number.
How often should the prompts be run?
Daily. AI engines return a different answer to the same question at different times, so a weekly check produces a number nobody can trust. Running daily gives seven readings per question per engine each week, and that weekly share is the smallest unit worth comparing. Agenzy reads four-week trends and reports them per engine.
What does prompt tracking cost?
Tracking tools are sold as a subscription, priced by how many questions you run and how many engines you read, from small entry plans up to enterprise contracts. Agenzy includes the tracking in the GEO service, which starts from 5,000 EUR per month ex VAT and covers the work behind the numbers as well as the measurement. We compare what agencies charge in what a GEO agency costs in 2026.
Which AI engines should I track?
Track ChatGPT first, then Google AI Overviews, then Gemini. Statcounter put ChatGPT at 70.9% of AI chatbot web traffic in the UK and 65.0% in the US in July 2026, with Gemini second among the chatbots. We still put Google AI Overviews ahead of Gemini, because they sit on Google search and reach a far larger audience. Browser data understates Gemini, since most of its use happens inside Google's apps.
Is prompt tracking the same as rank tracking?
No. Rank tracking records where your page sits in a list of Google links, and prompt tracking records whether an AI answer names your brand at all. The two are related: pages ranked first in Google are cited by ChatGPT far more often than pages past the twentieth result, yet a top position guarantees nothing. AEO, AI SEO and AIO name the work; prompt tracking measures it.
Can I do prompt tracking by hand in ChatGPT?
Yes for a first look, no for a trend. Asking five questions once in ChatGPT tells you whether your brand shows up at all, and that is a useful hour. A trend needs logged-out sessions, a run every day and a written record of every answer. Nobody keeps that up by hand.
About Agenzy
Agenzy is a GEO agency based in Vilnius, working with brands in the US, the UK and across Europe. We get brands named and recommended inside ChatGPT, Gemini, Google AI Overviews, Perplexity, Claude and Copilot. We are an official Peec AI partner. As of September 2026: 500,000+ AI chats analysed, 15,000+ prompts tracked, 150+ audits completed, 1,000,000+ EUR generated for clients by AI search. The studies behind the numbers on this page are published at agenzy.lt/research.




