Written by Emilis Zabilius, co-founder of Agenzy.
In May 2026 the robots.txt on outcraft.ai blocked four AI crawlers, GPTBot, ClaudeBot, CCBot and Google-Extended, and Outcraft AI was named in 5 of 100 tracked buyer prompts. Ninety days of generative engine optimization (GEO) later, the brand held 98% share of voice on the tracked prompt "best AI tools for e-commerce failed payment recovery" in the 14 days to 2026-08-31. The number started moving before a single new page existed, first in Google AI Overviews, and those four crawler rules were the first item on the fix list.
Outcraft AI sells AI voice, SMS and email agents that recover failed payments and abandoned carts for e-commerce and SaaS companies in the US market, and the Outcraft case study carries that result with its charts. Agenzy does not publish revenue estimates for clients, so the chain evidenced here runs from AI answers to citations to Google impressions and clicks, and stops there.
Across all 100 tracked prompts, AI visibility read 0.25% in the week of 2026-05-18 and 11.90% in the week of 2026-08-24, which is 48 times the baseline reading. That 48x sits on a very small base: the baseline week carried six mentions in total, so it measures a move away from near zero.
Where Outcraft AI's AI visibility started in May 2026
Outcraft AI started in May 2026 with four AI crawlers blocked in its robots.txt, no structured data on the homepage, and a mention in 5 of 100 tracked buyer prompts. A software company whose competitors appear in every AI answer about its category usually has a retrieval problem before it has a content problem, and this is the documented version of that condition.
The table below is the May 2026 baseline, one row per thing we read and the instrument that read it, all of it captured before any work started.
What we read, with its unit and date | The reading | The instrument |
|---|---|---|
AI crawler access in robots.txt | GPTBot, ClaudeBot, CCBot and Google-Extended blocked | the live robots.txt file |
Tracked buyer prompts with a mention | 5 of 100 | Peec AI, 3 engines, daily |
AI visibility, week of 2026-05-18 | 0.25%, meaning 6 mentions across 2,436 AI answers | Peec AI |
Citations of outcraft.ai in AI answers | 20 in the week of 2026-05-18 | Peec AI |
Organic Google visits a month | about 41, against 7,000 to 37,000 for the brands holding 8% to 23% AI visibility | Semrush |
Structured data on the homepage | none; no Organization, Product or FAQPage JSON-LD in the page source | live page source |
URLs in the sitemap | 26, of which 14 were blog posts | the site's own sitemap |
The 0.25% carries its denominator, which is 6 mentions across 2,436 AI answers, and the closing reading of 11.90% rests on 2,068 answers, so both endpoints sit on almost the same sample size. Share of voice is a brand's share of all the brand mentions on a tracked prompt, so 98% means almost every brand named in those answers was Outcraft. In those answers Outcraft was being compared with funded voice AI platforms and with software suites carrying thousands of ranking pages.
The GEO method, week by week: what shipped in each of the fifteen weeks
The order on the fix list was crawler access and structured data first, pages second, then an audit of what was live and a nine-article batch aimed at the prompts still at zero. The table below gives each of the fifteen weeks: what we found or shipped, how we verified it, and the visibility Peec AI recorded across all 100 tracked prompts in that week.
Week beginning (Monday) | What we found or shipped | How we verified it | AI visibility that week (% of AI answers naming Outcraft, 100 prompts, 3 engines, Peec AI) |
|---|---|---|---|
2026-05-18 | Baseline read on the 100-prompt set; robots.txt found blocking GPTBot, ClaudeBot, CCBot and Google-Extended | Baseline counted in the tracker: 6 mentions across 2,436 AI answers, 5 of 100 prompts | 0.25% |
2026-05-25 | Kickoff; the 100-prompt set agreed and frozen; technical pack handed over: the robots.txt rules, structured data for the site and markup for the existing blog posts | Agreed prompt list matched line by line against the tracker | 0.19% |
2026-06-01 | FAQ copy for the buying questions shipped on 2026-06-03 | Each question mapped to the tracked prompt clusters before it shipped | 0.49% |
2026-06-08 | No new pages published | No publication this week, so nothing to verify | 1.16% |
2026-06-15 | No new pages published; the engines read what was already there | No publication this week, so nothing to verify | 2.52% |
2026-06-22 | Month-two content plan built on the tracked prompts; no new pages published | No publication this week, so nothing to verify | 1.90% |
2026-06-29 | Four articles in progress, each written for a group of tracked prompts | Nothing published yet, so nothing to verify | 3.39% |
2026-07-06 | Four articles live on outcraft.ai on 2026-07-10 with Article and FAQPage markup; metadata pass on three YouTube videos | Markup read live on all four on 2026-07-31 | 6.46% |
2026-07-13 | Four LinkedIn articles published on Outcraft's own accounts | Publication confirmed by Outcraft; the date is not in our own log | 9.36% |
2026-07-20 | Nothing shipped and nothing found; the week carries no entry in our log | No deliverable and no finding logged that week. | 5.25% |
2026-07-27 | Technical pack confirmed live: crawler rules removed, markup on the four July articles; live blog corpus audited, 41 of the 100 prompts still at 0% in the July read, zero-coverage prompt clusters mapped | Live robots.txt and page source read on 2026-07-31; tracker retrievals read for every live post | 4.26% |
2026-08-03 | The fall diagnosed as engine-side and reported with the Search Console cross-check; fix list for the split apex and www hosts handed over on 2026-08-05 | Gemini grounding rate read at 97% then 65%; Search Console checked in the same days | 4.67% |
2026-08-10 | Domain fixes followed up; no new pages published | No publication this week, so nothing to verify | 4.54% |
2026-08-17 | Domain consolidation live; nine articles drafted for the prompt clusters still at 0% | Read live on 2026-08-20: apex returning 301, sitemap and schema ids on the www host; every draft checked against its tracked prompt | 5.25% |
2026-08-24 | The first four of the nine live on 2026-08-25 | All four checked live on 2026-08-25 with their markup in place | 11.90% |
The number moved in three stretches in those fifteen weeks: Google AI Overviews picking up the existing pages in June, and the two publication weeks in July and August, with five flat weeks between them.

The visibility column has two stretches with no publication behind them, and they went opposite ways. Weeks one to five contain no new page at all, and the reading still went from 0.25% to 2.52%. Most of that came from Google AI Overviews, which reads pages through Google Search's own crawler and was never behind the blocked rules. Retrieval moved with it: outcraft.ai was retrieved as a source in 4.4% of AI answers by the week of 2026-06-15, from 0.4% in the baseline week.
The week of 2026-08-17 holds nine drafts and no publication, and it stayed inside the flat band at 5.25%.
The lag between publishing and movement is not fixed here. The four July articles went live on 2026-07-10 and the strongest week of that phase was 2026-07-13, at 9.36%. Within three weeks of that publication the tracked SaaS free-trial prompt went from 0% to 23.8% and the mobile checkout abandonment prompt from 4.4% to 28.6%. The August batch went live on 2026-08-25 and the same week closed at 11.90%, with 616 citations of outcraft.ai recorded in it.
What happened in the weeks after 2026-08-24
The engagement carried on after the case study window closed, and the three complete weeks that followed test whether one week carried the result. Each cell is the share of that engine's answers to the same 100 tracked prompts that named Outcraft, as Peec AI recorded it.
Week beginning (Monday) | AI visibility, all three engines (% of AI answers naming Outcraft, 100 tracked prompts) | ChatGPT | Gemini | Google AI Overview |
|---|---|---|---|---|
2026-08-17 | 5.25% | 1.86% | 9.14% | 4.70% |
2026-08-24 | 11.90% | 3.00% | 19.00% | 13.77% |
2026-08-31 | 15.10% | 4.14% | 20.71% | 20.59% |
2026-09-07 | 11.26% | 4.14% | 14.29% | 15.46% |
Google AI Overview and ChatGPT ended these three weeks above the level they closed the case study at, and Gemini gave part of its August spike back.

Across the 14 days to 2026-09-15, visibility on the whole set reads 12.25%, from 509 mentions in 4,154 AI answers, and 57 of the 100 tracked prompts carried a mention. The prompt "best AI tools for e-commerce failed payment recovery" reads 42.86% visibility and 90.91% share of voice in that window, with three larger brands now appearing on it where the case study window had one.
So the level of 2026-08-24 was beaten the following week and sat just below it the week after, with Google AI Overview carrying the gain while Gemini gave part of its spike back.
The five weeks when the line went sideways
Between 2026-07-20 and 2026-08-23 Outcraft's AI visibility sat between 4.26% and 5.25% for five straight weeks, after a mid-July peak of 9.36%.
Two engine-side observations sit around that fall, both made on 2026-08-04 with the client's Search Console export open beside the tracker. The share of Gemini answers grounded in live search fell from 97% to 65% in our reading. In Google AI Overviews, the answers on the prompts Outcraft lost moved to competitor listicles, while the engine kept retrieving outcraft.ai at about the same rate.
Attribution in either direction is a reading. We cannot prove the fall was engine-side any more than we can prove the July rise came from the articles, because the tracker records what the engines said and not why they said it.
The cross-check is the measured part. Google organic had not fallen at all: impressions were at a three-month high and the average position had improved, so the decline showed only in the AI-layer readings, which locates it without explaining it. The corpus audit, the cluster map and the domain consolidation fixes all happened inside those five weeks.
The monthly report is where that cross-check reaches the client, and what a monthly report should show sets out how we build one.
The prompt set was agreed before any work started
The 100 tracked prompts were agreed in writing with Outcraft before the work began and frozen for the engagement. A prompt set is the measurement, so a set that can be edited mid-flight can be made to show anything.
The set was built from the pitch-stage tracking and a gap analysis of the category. Useful prompts from the pitch set were kept, how-to questions were rewritten as buyer requests that send an engine to the web, and new prompts covered the gaps.
The composition is 50 B2B SaaS prompts and 50 e-commerce prompts, run daily against ChatGPT, Gemini and Google AI Overview. Not one of them names Outcraft, because a prompt carrying the brand name measures recall of that name, and the set is there to show whom the engines recommend when a buyer describes a problem. Ten prompts name another vendor, so the list published with the case study is the remaining 90.
Freezing the set has a cost we accepted. When a topic is worth writing and no tracked prompt sits under it, the piece still gets written and it will not show up in the measurement. Adding the prompt that a new page happens to answer moves the number without changing a single AI answer.
The set is also what every deliverable is checked against. An article gets scheduled when a named prompt in the set sits at 0% and the engines are answering it with somebody else's page. That is why the August batch was aimed at the clusters the July read showed at zero: inbound speed and routing, alternatives, churn, integrations.
The pages and off-site work that earned the citations
The citations came from Outcraft's own pages. Across the window outcraft.ai was cited 3,000+ times in AI answers, from 20 in the baseline week to 616 in the week of 2026-08-24. By that week the engines retrieved a page from the domain in 20.2% of all answers they gave to the tracked questions. The table below lists the work produced in the 90 days, with its count, its date and where it was published.
Type of work | What shipped in the 90 days, with dates | Where it was published |
|---|---|---|
Articles written for tracked prompts | 4 in month two, live 2026-07-10; 9 drafted in month three with the first 4 live on 2026-08-25 | outcraft.ai |
Structured data | company, website, product and review blocks on the homepage, then Article and FAQPage on every new article | outcraft.ai |
FAQ set covering the buying questions the tracked prompts ask | delivered in month one | Outcraft's FAQ page |
YouTube | metadata, chapters and captions rebuilt on 3 existing videos | Outcraft's own channel |
LinkedIn articles | 4, in the week of 2026-07-13 | Outcraft's own accounts |
Third-party placements | outreach prepared, none landed inside the 90 days | nothing published |
Every line in that table is owned media, published on properties Outcraft controls.
Listicles and review pages are a large share of what AI engines cite in this category. No placement landed inside the window, so every result reported for it rests on owned media plus the technical work that made those pages readable. A program that depends on earned placements runs a slower curve than this one did.
Every article in the batch followed the same format. Each piece answers one buyer question in its first two lines, carries that question as its heading, and states every number with its window attached. AI engines score passages against the exact question a buyer typed, which is the principle behind our guide to how to rank on ChatGPT.
The tools that ran the work and checked it
Peec AI is the tracker, and we are an official Peec AI partner. It ran the 100 prompts daily across ChatGPT, Gemini and Google AI Overview, and 32,000+ AI answers were analyzed through 2026-08-31. Every visibility, citation and share-of-voice figure here comes from that instrument.
Around it we run a daily prompt-level capture that keeps the per-prompt history in our own store, and a publishing pipeline that checks every draft against its tracked prompt. A bot-access check requests live pages with each AI crawler's user agent, and on 2026-08-20 it returned 200 with no challenge for GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot.
What changed in the engines during the 90 days
Each engine picked Outcraft up on its own schedule. Google AI Overview registered its first mention of Outcraft in the week of 2026-05-25 at 0.15%, then 0.46%, then 2.25% in the week of 2026-06-08. Gemini did not name the brand at all until the week of 2026-06-29, when it read 2.86%, a true zero before that.
By the 14 days to 2026-08-31, Gemini was the strongest engine at 15.0% and Google AI Overview sat at 10.5%. That window covers the weeks beginning 2026-08-17 and 2026-08-24, so it does not include the week beginning 2026-08-31 in the weekly engine table.
The two late-July changes on the engines' side, Gemini's lower grounding rate and the source shift in AI Overviews, show up in the weekly series more clearly than anything shipped that month. Neither was predictable from anything on the client's site.
Where it did not work: ChatGPT
ChatGPT is the engine where this engagement did least. Outcraft read 2.4% there in the 14 days to 2026-08-31, against 15.0% on Gemini in the same window. It has since risen by about 70% off that base, to 4.14% in the week of 2026-09-07 and 4.07% across the 14 days to 2026-09-15, and in that last window it was still the weakest of the three engines.
Our reading is that ChatGPT leans on third-party lists and review pages in this category, which is the layer that did not convert inside the 90 days. A single averaged visibility figure across the three engines would have hidden the gap, which is why the reporting Outcraft receives splits the figures by engine.
The results on the tracked prompts
The target agreed before the work started was 20 tracked prompts with a mention in the final two weeks. The count was 31 in the window of 2026-08-10 to 2026-08-23, which is the reading reported for that window. By 2026-08-31 it was 51 of 100, and in the 14 days to 2026-09-15 it was 57 of 100, against the 5 the engagement started with.
The five category prompts below come from the same tracked set and are read over the 14 days to 2026-08-31, one row per prompt, with both metrics as Peec AI recorded them.
Tracked prompt | Outcraft visibility (% of AI answers naming Outcraft, 14 days to 2026-08-31) | Outcraft share of voice (% of all brand mentions on that prompt, same window) | Other tracked brands named |
|---|---|---|---|
Best AI tools for e-commerce failed payment recovery | 59.5% | 97.8% | one other tracked brand at 2.2% of mentions |
What AI tool can recover failed payments and prevent churn at the same time | 33.3% | 100% | none |
Best AI tools for SaaS failed payment recovery automation | 22.0% | 85.0% | two software suites, 15.0% of mentions between them |
Best AI tools for automated Stripe failed payment recovery | 22.0% | 100% | none |
Best AI tools for recovering abandoned Stripe subscriptions | 21.4% | 100% | none |
The rows at 100% mean no other tracked brand was named on that prompt at all. Visibility is the share of answers naming the brand, which is why a prompt reads 22.0% and 100% at once: most answers named nobody, and every answer that named somebody named Outcraft. The e-commerce failed payment recovery prompt reads 97.8%, which rounds to the 98% on the case study.
Can visibility in AI answers turn into revenue
On the Outcraft AI engagement that Agenzy ran, AI visibility turned into measurable search demand, and the evidence stops one step short of revenue. Comparing matching 14-day windows, 2026-05-28 to 2026-06-10 against 2026-08-14 to 2026-08-27, Google impressions went from 9,896 to 88,987, clicks from 143 to 551, and the impression-weighted average position from 12.4 to 9.1. By August 2026, Search Console was logging full conversational queries landing on pages built for the AI engines, which fits people checking in Google what an assistant told them.
Google traffic and AI visibility read different signals. Outcraft's organic traffic sat two orders of magnitude below the brands it was compared with inside AI answers, and by 2026-08-31 eight tracked competitor brands that had sat above it in May 2026 sat below it.
How to read a GEO case study, including this one
Four questions settle whether a GEO case study is evidence or a screenshot. A demand generation lead shortlisting AI search agencies can ask them of any case page, including this one, and how to choose a GEO agency covers the rest of that shortlist.
What was the number before the work started, in absolute values as well as percentages? Here it was 0.25%, which is 6 mentions across 2,436 AI answers, and 5 of 100 tracked prompts.
Which tool measured it, and across which dates? Peec AI, daily, on ChatGPT, Gemini and Google AI Overview, from the week of 2026-05-18 to the week of 2026-09-07.
What shipped in which week? The method table names the deliverable, the verification and the reading for each of the fifteen weeks.
Which weeks went nowhere? Five of them, 2026-07-20 to 2026-08-23, and one of those five carries no deliverable at all.
What this B2B SaaS GEO case study does not prove
This is one engagement with one client in one thin competitive category, measured over 90 days.
There was no holdout on this engagement. No matched set of prompts was held back to show what would have happened without the work, so the timeline establishes a sequence in time. Google's May 2026 core update also rolled out between 2026-05-21 and 2026-06-02, over the first three weeks of the series, so the early movement in AI Overviews cannot be separated from it. We run holdouts on interventions now, and this case predates that practice.
The Google multiple comes off a very small base. Impressions running 9x sounds larger than it is when the starting point is 707 impressions a day.
Part of the AI visibility gain is the site becoming readable. A brand whose robots.txt already allows the AI crawlers has banked that part and should expect a smaller jump from a higher base.
Who this would not work for
This method has been tested once, on a company whose category prompts were still thin, that could publish pages, and whose product pages an engine could read. In three situations we would expect a much weaker result, and would say so before a contract.
A category already owned by six listicles
Where the engines answer a buyer question by quoting the same handful of third-party lists and review sites every time, owned pages alone will not displace them. Outcraft's own numbers show the shape of it: 100% share of voice on prompts where almost nobody was named, and its weakest reading on ChatGPT, the engine where those third-party sources appear to weigh most.
Winning a category like that means earning placements on the pages the engines already quote, which is slower, more expensive and less certain than anything in this engagement's timeline.
A company that cannot ship pages
Both of the large jumps in this case followed a publication week. The August batch went live in one week. A brand with a six-week legal review on every article, or nobody with the authority to approve copy, gets a timeline set by its own approval queue.
A brand with no product page an engine can read
If the product lives behind a login, inside a PDF or in a video with no transcript, there is nothing for an engine to retrieve. Where those pages do not exist yet, the first months go into building them and the visibility curve starts later than this one did.
In two more cases the right answer is to wait. A brand whose buyers do not ask assistants before they buy has nothing to gain yet, and one question to existing customers settles it. A company measuring nothing today needs a baseline first, since a later percentage can only be checked against a number read before the work. What the first 90 days of an engagement contain is covered in GEO for B2B SaaS, the first 90 days.
FAQ
What is a GEO case study?
A generative engine optimization case study documents how a brand's presence in AI answers changed over a named period, with the measurement tool, the prompt set and the dates attached. A useful one shows the value before the work as well as after it, and the sequence of what was shipped. Without a starting number, a percentage in a GEO case study cannot be checked by the reader.
How long did it take before anything moved for Outcraft AI?
The first movement came in week three, before any new page existed: visibility went from 0.25% in the week of 2026-05-18 to 2.52% in the week of 2026-06-15, most of it in Google AI Overviews. Gemini named the brand for the first time in the week of 2026-06-29. The large gains needed published articles and arrived in month two and month three.
What was fixed first?
The crawler rules were first on the fix list handed over in May 2026. The robots.txt denied GPTBot, ClaudeBot, CCBot and Google-Extended, which kept the site out of model training and out of Gemini's grounding, while ChatGPT search and Google AI Overviews fetch pages with other crawlers. Structured data came in the same handover, content followed from month two, and off-site outreach came last.
Does blocking AI crawlers really stop you appearing in ChatGPT?
It depends on which crawler the file blocks. ChatGPT search fetches pages with OAI-SearchBot and ChatGPT-User, while GPTBot collects training data, so a file that blocks only GPTBot keeps a site out of future training without removing it from ChatGPT's search answers. Google-Extended works the same way for Gemini: it decides whether Gemini may use a site for training and grounding, and it does not affect Google Search or AI Overviews. Outcraft's file blocked the training and grounding family, and Gemini was the last engine to name the brand.
How many prompts were tracked, and who chose them?
One hundred, split evenly between B2B SaaS and e-commerce questions, agreed in writing with Outcraft before the work began and frozen for the engagement. None of them names Outcraft, and because the set never changed, every weekly reading measures the same questions. Ten prompts name another vendor and sit outside the published list of 90.
Did the gain hold after the 90-day window closed?
It held, with one week slightly below the case-study level: the week of 2026-08-24 closed at 11.90%, the week of 2026-08-31 at 15.10% and the week of 2026-09-07 at 11.26%. Across the 14 days to 2026-09-15, visibility on the tracked set reads 12.25% and 57 of the 100 prompts carried a mention. Google AI Overview held the largest share of the gain, Gemini returned part of its August spike, and ChatGPT improved from the lowest base of the three.
Would this work for a larger company?
Size is rarely what decides it. Blocked crawlers, five prompts with a mention and no structured data describe a condition, and companies of every size sit in it. The category is what changes the answer. Where third-party lists already hold the answers to the buying questions, the work moves to earning placements on those lists, and that runs slower than this engagement's timeline.
About Agenzy
Agenzy is a GEO and AEO agency: we get brands named and recommended inside ChatGPT, Gemini, Google AI Overviews, Perplexity, Claude and Copilot. We are an official Peec AI partner with 500,000+ AI chats analysed, 15,000+ prompts tracked and 150+ audits completed, and we count 1,000,000+ EUR generated for clients by AI search from clients' own CRMs and checkouts, where the buyer said an AI answer sent them.
Outcraft AI is our client, and each figure here comes from our tracking of the engagement or from the instrument named beside it.




