An AI engine answers a buyer's question by retrieving pages and citing some of them, and the pages it cites decide which brands appear in the answer. We counted the cited sources behind 394,180 AI answers to buyer questions about 130 European businesses, three quarters of them in Lithuania: 104 of the 138 tracked market projects ran from Lithuania and 34 from other markets. Each source was sorted by engine, by source type, and by whether the page belonged to the brand, to a named competitor or to a third party.
The same pattern holds in every cut. Citations are spread across a large number of websites, the three main engines pick different ones for the same question, and a brand's own website is a small slice of what gets cited about its category. The first study in this series, The State of AI Search Visibility in Europe 2026, measured how often brands appear in AI answers, and this one measures where those answers come from.
Key numbers (TL;DR)
No single website holds more than 2.1% of the 2,283,923 AI citations in Agenzy's 2026 source study, and it takes 1,057 different websites to cover half of everything the AI engines cited (n=138 business projects).
ChatGPT, Google AI Overviews and Perplexity share only 2 of their top 10 cited sources when answering the same questions about the same brand, according to Agenzy's 2026 source study (n=109 business projects).
Google AI Overviews takes 10.8% of its citations from user-generated content against Perplexity's 3.8%, while Perplexity takes 13.6% from news and media against AI Overviews' 5.2%, Agenzy's 2026 study of 2,283,923 AI citations found (n=137 and 133 business projects).
YouTube earns 4.7% of Google AI Overviews citations and 0.03% of ChatGPT citations in Agenzy's 2026 source study of 2,283,923 AI citations (n=137 and 138 business projects).
Reddit is the single most cited domain in Agenzy's 2026 AI source study, taking 2.1% of 2,283,923 citations and appearing in 92.8% of 138 business projects, with ChatGPT citing it 5.0x more heavily than Perplexity.
Brands' own websites earned 3.7% of the 2,283,923 AI citations in Agenzy's 2026 study while their named competitors earned 16.9%, a 4.6x gap (n=138 business projects).
34.8% of the businesses in Agenzy's 2026 source study did not have their own website among the 30 sources AI engines cited most for their own category (n=138 business projects).
Methodology
Agenzy analysed 2,283,923 AI citations from 394,180 AI answers across 138 tracked market projects covering 130 European businesses in 29 countries, using ChatGPT, Google AI Overviews, Gemini and Perplexity.
The instrument was Peec AI, which runs a fixed set of buyer prompts on each engine every day, stores the answer and records every source behind it. Peec's definitions apply throughout: sources are all URLs the AI model accessed while generating the response, and "citations are sources explicitly referenced in the response text." A retrieval in this paper is a page the engine pulled for an answer, whether or not it cited that page, and Peec records both counts for every URL.
A domain is the hostname Peec stores for a cited URL without the www prefix, so rekvizitai.vz.lt and vz.lt are separate rows, and the Wikipedia figure later in the paper pools every language edition.
Source types are Peec's classification of each domain: Corporate, Competitor, Editorial, Institutional, UGC (user-generated content), Reference, and "You" for the tracked brand's own domain. "Competitor" means the rivals each project itself configured. One project's custom source class (0.2% of all citations) is left out of the tables. No tracked business is named in this paper and no figure identifies an individual project.
Dataset | Value |
|---|---|
Tracked market projects | 138 |
Unique businesses | 130 |
Countries prompts ran from | 29 |
Tracked prompts | 11,296 |
AI answers | 394,180 |
Citations, domain table | 2,283,923 |
Citations, URL table | 2,037,717 |
Classified URL rows | 333,396 |
Domains cited | at least 100,400 |
Window | 2025-01-01 to 2026-09-07 |
Peec reports citations in two tables. The domain table holds 2,283,923 citations and is the base for every citation share in this paper. The URL table holds 2,037,717 citations across 333,396 classified rows and is the base for the page-type figures.
The window covers everything each project ever ran. A project tracks one business in one market, so a business tracked in two countries counts once as a business and twice as a market project. ChatGPT ran in 138 projects, Google AI Overviews in 137, Gemini in 16 of the source tables (18 in the answer counts) and Perplexity in 133. Every Gemini figure carries its own n, and Gemini sits outside the three-engine comparisons. Google AI Mode and Claude appeared in one project each and are excluded from every per-engine cut, and Agenzy's own tracking project is excluded from every figure.
The sample is Agenzy's client and pitch dataset, so it reflects which businesses approached a GEO agency in 2025 and 2026. The categories are the ones those businesses asked to have tracked. The figures are a measurement of one agency's book, and the Limitations section returns to this.
How concentrated are the sources AI engines cite?
No single website holds more than 2.1% of the 2,283,923 AI citations in Agenzy's 2026 source study, and it takes 1,057 different websites to cover half of everything the AI engines cited (n=138 business projects).
The top 10 domains hold 8.7% of all citations, the top 50 hold 16.8% and the top 100 hold 21.9%, out of at least 100,400 domains cited. Any published list of "the websites AI cites most" therefore describes a thin layer at the top of a very long tail. The single most cited domain, reddit.com, sits at 2.1%.
Cut of the dataset | Share of 2,283,923 citations |
|---|---|
Most cited domain | 2.1% |
Top 10 domains | 8.7% |
Top 50 domains | 16.8% |
Top 100 domains | 21.9% |
Domains needed to reach 50% | 1,057 |

Figure 1. Cumulative share of the 2,283,923 citations by domain rank: the top 100 domains reach 21.9% and the curve crosses 50% at 1,057 domains (n=138 projects).
Inside a single project the picture inverts. The top 10 domains hold a median 35.3% of that project's citations, and a median of 23 domains covers half of them. Lithuanian-market projects concentrate slightly more than the rest, with a median top-10 share of 36.3% against 30.3% in other markets (n=104 and 34). Each project asks about one category in one market, and each engine retrieves from the pages that already answer that category's questions in that market's language, which is why the pooled table fragments while each project's own table stays concentrated.
Do ChatGPT, Google AI Overviews and Perplexity cite the same sources?
ChatGPT, Google AI Overviews and Perplexity share only 2 of their top 10 cited sources when answering the same questions about the same brand, according to Agenzy's 2026 source study (n=109 business projects).
The pairwise overlap is small as well. For the same brand and the same prompts, ChatGPT and Google AI Overviews share a median of 4 of their top 10 cited domains, and in 5 of 119 projects that pair shared none. ChatGPT and Perplexity share 3, and Google AI Overviews and Perplexity share 4. A brand cited by one engine has no assurance of being cited by another, which is why every later finding in this paper reports each engine on its own.
Engine pair | Median shared domains in the top 10 | Projects (n) |
|---|---|---|
ChatGPT and Google AI Overviews | 4 of 10 | 119 |
ChatGPT and Perplexity | 3 of 10 | 128 |
Google AI Overviews and Perplexity | 4 of 10 | 109 |
All three engines | 2 of 10 | 109 |
In 14.7% of projects (16 of 109) no single domain appeared in all three top-10 lists. Widening the cut to the top 25 sources lifts the three-engine median to 6 of 25 (n=103), so the disagreement holds at depth as well as at the top. The comparison only includes projects where all three engines had at least 10 cited domains, which leaves out the smallest projects.

Figure 2. Median shared domains in the engines' top-10 source lists for the same brand: 4, 3 and 4 for the three pairs and 2 for all three engines together (n=119, 128, 109 and 109 projects).
Why the engines disagree
The engines run on different indexes and different selection layers. ChatGPT retrieves from OpenAI's own index and still pulls from Google, Microsoft and other providers alongside it, as Peec AI documented from ChatGPT's server-side events in September 2026. Google states that "the best practices for SEO remain relevant for AI features in Google Search (such as AI Overviews and AI Mode)". The same page says both surfaces may use a "query fan-out" technique, issuing multiple related searches across subtopics and data sources to build one response (Google's AI features guide). Perplexity crawls with its own bot and builds its own candidate pool. Different candidate pools go through different rerankers, and each answer cites what came through that engine's pipeline.
Which types of sources does each AI engine prefer?
Google AI Overviews takes 10.8% of its citations from user-generated content against Perplexity's 3.8%, while Perplexity takes 13.6% from news and media against AI Overviews' 5.2%, Agenzy's 2026 study of 2,283,923 AI citations found (n=137 and 133 business projects).
The two n values in that sentence are the 137 projects that ran on Google AI Overviews and the 133 on Perplexity; the table below carries each engine's own base.
Source type | ChatGPT | Google AI Overviews | Gemini (n=16) | Perplexity | All engines |
|---|---|---|---|---|---|
Corporate | 46.7% | 46.5% | 65.4% | 47.2% | 48.0% |
Competitor | 17.4% | 18.8% | 14.0% | 13.2% | 16.9% |
Institutional | 11.9% | 5.5% | 0.9% | 8.5% | 8.6% |
Editorial | 7.7% | 5.2% | 5.1% | 13.6% | 7.6% |
UGC | 4.8% | 10.8% | 5.7% | 3.8% | 6.6% |
Reference | 4.0% | 3.3% | 2.1% | 5.5% | 3.9% |
You (own domain) | 3.3% | 4.5% | 4.0% | 3.2% | 3.7% |
Other | 3.9% | 5.2% | 2.7% | 4.9% | 4.5% |
Projects (n) | 138 | 137 | 16 | 133 | 138 |
Citations | 1,045,566 | 731,699 | 110,394 | 343,055 | 2,283,923 |

Figure 3. Share of each engine's citations by Peec source type, with UGC running from 10.8% on Google AI Overviews to 3.8% on Perplexity and editorial from 13.6% on Perplexity to 5.2% on AI Overviews (n=138, 137, 16 and 133 projects).
Institutional pages are ChatGPT's distinctive habit. It sends 11.9% of its citations to government, health and EU domains, against 8.5% on Perplexity and 5.5% on Google AI Overviews. ChatGPT's five most cited domains include the Lithuanian central bank lb.lt at 1.0% of its citations, the tax authority vmi.lt at 0.8% and europa.eu at 0.8%.
An equal-depth cut of the top 50 domains per engine per project keeps the same UGC order, 6.0% on ChatGPT (n=126), 14.6% on Google AI Overviews (n=124) and 4.2% on Perplexity (n=118), so table depth is not driving the result.
Where each engine's source mix comes from
Our reading is that source personality follows the product each engine is built on. Google AI Overviews sit on top of Google web results, where forum threads and videos rank for many buyer questions, and the overview inherits that mix. Perplexity built its product around visibly cited answers, and here those citations lean editorial. ChatGPT reaches official pages heavily in this dataset, which fits prompts about regulated categories (finance, health, public procurement) in a market where the authoritative page is often a government one. None of the three vendors documents its source selection, so the mechanism is inferred from the output.
Does Google AI Overviews cite YouTube and Facebook more than ChatGPT?
YouTube earns 4.7% of Google AI Overviews citations and 0.03% of ChatGPT citations in Agenzy's 2026 source study of 2,283,923 AI citations (n=137 and 138 business projects).
YouTube took 34,544 Google AI Overviews citations and appeared in 72.3% of the 137 projects on that engine. On ChatGPT it took 292 citations across 21.7% of 138 projects, and on Perplexity 0.6% of citations across 46.6% of 133 projects. Facebook follows the same shape: 1.9% of AI Overviews citations (14,048), against 51 citations in total on ChatGPT, out of 1,045,566. The gap is two orders of magnitude, and the rows missing from the capped tables, domains cited once or not at all, cannot close it.

Figure 4. YouTube and Facebook as a share of each engine's citations: 4.7% and 1.9% on Google AI Overviews against 0.03% and 51 citations on ChatGPT (n=137 and 138 projects).
Across the whole dataset youtube.com is the second most cited domain, at 1.6% of all citations and present in 88.4% of projects, and facebook.com appears in 94.2% of projects at 0.7%. Both positions come from Google AI Overviews.
Which website do AI engines cite the most?
Reddit is the single most cited domain in Agenzy's 2026 AI source study, taking 2.1% of 2,283,923 citations and appearing in 92.8% of 138 business projects, with ChatGPT citing it 5.0x more heavily than Perplexity.
Reddit's 48,018 citations split unevenly by engine: 3.0% of ChatGPT's citations, 1.7% of Google AI Overviews' and 0.6% of Perplexity's. The project shares follow the same order, with Reddit present in 77.5% of ChatGPT projects, 61.3% of Google AI Overviews projects and 46.6% of Perplexity projects (n=138, 137 and 133).
Reddit alone is 32.0% of the 149,900 citations Peec classed as user-generated content, and together with YouTube, Facebook, Instagram and LinkedIn it accounts for 72.1% of them. Facebook sits inside that count with 11,769 of its 15,363 citations, the rows Peec classed as UGC. Its remaining rows are classed Other, which is why the table below lists it as Other.

Figure 5. Reddit's share of each engine's citations: 3.0% on ChatGPT, 1.7% on Google AI Overviews and 0.6% on Perplexity (n=138, 137 and 133 projects).
Rank | Domain | Peec class | Citations | Share | Projects |
|---|---|---|---|---|---|
1 | reddit.com | UGC | 48,018 | 2.1% | 128 (92.8%) |
2 | youtube.com | UGC | 37,264 | 1.6% | 122 (88.4%) |
3 | lrv.lt | Other | 17,908 | 0.8% | 80 (58.0%) |
4 | vz.lt | Editorial | 16,868 | 0.7% | 102 (73.9%) |
5 | facebook.com | Other | 15,363 | 0.7% | 130 (94.2%) |
6 | lb.lt | Institutional | 12,983 | 0.6% | 25 (18.1%) |
7 | google.com | Other | 12,894 | 0.6% | 83 (60.1%) |
8 | paslaugos.lt | Reference | 12,752 | 0.6% | 80 (58.0%) |
9 | vmi.lt | Institutional | 12,331 | 0.5% | 34 (24.6%) |
10 | europa.eu | Institutional | 12,147 | 0.5% | 73 (52.9%) |
The top 10 shows the shape of the sample. Reddit and YouTube lead, then come the Lithuanian government portal, the Lithuanian business daily, the central bank and tax authority, and the EU. Wikipedia holds 0.09% of citations here while appearing in 63.0% of projects.
Ahrefs' September 2026 count of the domains ChatGPT cites most puts Reddit at 16.8% of ChatGPT citations and Wikipedia at 7.0%, measured on US queries. In this dataset Reddit is 3.0% of ChatGPT citations and Wikipedia 0.13% (1,329 of 1,045,566), measured on buyer questions run mostly from Lithuania. The two counts use different prompt sets, different markets and different windows, so they are not a like-for-like comparison.
The market cut inside this dataset points the same way. Lithuanian-market projects give Reddit 1.4% of citations and user-generated content 4.1%, while projects run from other markets give Reddit 2.9% and UGC 9.5% (n=104 and 34).
How often do AI engines cite a brand's own website compared with its competitors?
Brands' own websites earned 3.7% of the 2,283,923 AI citations in Agenzy's 2026 study while their named competitors earned 16.9%, a 4.6x gap (n=138 business projects).
The pooled ratio understates the typical case. Per project, the median own share is 1.6% and the median competitor share is 15.4%, a median gap of 7.8x. Company websites of every kind (own, competitor and other corporate) take 68.6% of all citations, so company pages are the bulk of what the engines cite. The engines cite other companies' pages far more often than the tracked brand's.

Figure 6. Share of the 2,283,923 citations going to the brand's own website and to its named competitors, 3.7% against 16.9%, with the per-project medians of 1.6% and 15.4% alongside (n=138 projects).
Among the 125 projects that earned at least one own-domain citation, the median own share is 2.0% and the highest is 17.5%. Part of the pooled gap is arithmetic, since each project's competitor set holds several rival domains against one own domain. A buyer's category question ("best accounting software for a small company") is answered from listicles, directories and rivals' comparison pages, and the tracked brand's site is retrieved only when it carries a page that answers that question.
How many businesses are missing from the sources AI engines cite about their category?
34.8% of the businesses in Agenzy's 2026 source study did not have their own website among the 30 sources AI engines cited most for their own category (n=138 business projects).
The own domain appears among the project's top 30 cited sources in 65.2% of projects (90 of 138). Across the full tables the own domain is cited somewhere in 90.6% of projects, at a median rank of 10, holding a median 2.0% of the project's citations. The typical brand is therefore cited for its own category, with its website around tenth among the sources the engines used to answer questions about it.
Engine | Projects where the own domain is cited | Median rank when cited | Median own share when cited | Projects with an own-domain citation (n) |
|---|---|---|---|---|
ChatGPT | 85.5% | 12 | 1.7% | 118 of 138 |
Google AI Overviews | 67.2% | 6 | 3.1% | 92 of 137 |
Perplexity | 66.9% | 8 | 2.3% | 89 of 133 |

Figure 7. Share of projects where each engine cited the brand's own domain, and its median rank when cited: 85.5% at rank 12 on ChatGPT, 67.2% at rank 6 on Google AI Overviews, 66.9% at rank 8 on Perplexity (n=138, 137 and 133).
The per-engine split runs in opposite directions. ChatGPT cites the brand's own site in the most projects (85.5%) and ranks it lowest when it does (median 12). Google AI Overviews and Perplexity cite it in about two thirds of projects and rank it higher (median 6 and 8). Where the site is cited, its median share of that engine's citations is 1.7% on ChatGPT, 3.1% on Google AI Overviews and 2.3% on Perplexity (n=118, 92 and 89). Own domains are identified by Peec's "You" class, which requires the domain to be registered on the brand. 129 of the 138 projects have an own domain Peec can classify, and the remaining 9 are the ceiling of that error. 125 of the 129 earned at least one own-domain citation, which is the 90.6% of all 138 projects quoted above, so 4 classifiable own domains were never cited.
Why the split runs in opposite directions
ChatGPT produced the largest citation table of the three engines, 1,045,566 citations against 731,699 for Google AI Overviews and 343,055 for Perplexity. More domains make it into a ChatGPT answer as a result, and the brand site gets in more often and further down the list. Perplexity is the most selective at the citation step, pulling 2.24 pages for every one it cites against 1.18 for Google AI Overviews and 1.36 for ChatGPT (n=132, 136 and 138). On the two selective engines a brand site appears only when one of its pages answers the question, and then it sits near the top.
What this means for businesses
Getting cited is mostly work on other people's pages, and it differs by engine: a brand's own website earned 3.7% of the citations in this dataset, and the three engines share 2 of their top 10 sources. Crawler access in robots.txt comes first, because OpenAI's crawler documentation states that "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links", and Google's AI surfaces read what Googlebot reads. The server and CDN layer comes next, since an Allow line in robots.txt means nothing when a firewall rule blocks the same agent, and Cloudflare's AI Crawl Control panel shows which AI services reach a site and which are blocked. After access, add schema that matches the visible page, then put the answer in the first lines of every page that targets a buyer question, the homepage included, since homepages absorbed 17.5% of the 2,037,717 URL-level citations here. Then work the pages each engine trusts for the category: YouTube for Google AI Overviews (4.7% of its citations), Reddit for ChatGPT (3.0%) and editorial pages for Perplexity (13.6%), with the category's own list built by asking the buyer's questions on each engine and recording every domain cited. The step-by-step version of this sequence is in how to rank on ChatGPT in 2026. Agenzy's GEO service runs it for clients, with the citation audit done before any content is written.
Limitations
The sample is an agency's client and pitch dataset: the businesses that approached Agenzy or that Agenzy audited for a pitch between 2025 and 2026. It therefore over-represents Lithuania (104 of 138 projects) and the categories where small and mid-sized European businesses buy GEO work. The 34 projects run from other markets are spread across the remaining countries in the dataset, so that cut is a contrast group and describes no single market. Nothing here is a random sample of the European web, and every figure should be quoted as a measurement of this dataset.
Source classification belongs to the tool. Peec assigns each domain a class per report row, and 3,011 of the at least 104,556 distinct domains seen across the overall and per-engine tables (2.9%) carry more than one class. The at least 100,400 figure quoted earlier counts the overall tables alone, and both counts are floors. Individual domains are sometimes misclassified (facebook.com is "Other" in the overall table and UGC in some per-engine rows), so class shares carry a margin of a few percentage points and no single domain's class is treated as a finding.
Table depth varies by project. Rows were pulled top-N by citation count, between 50 and 5,329 rows per project (median 1,000), and most projects hit that cap on at least one table, so the missing rows are domains cited once or not at all. Shares, per-engine comparisons and concentration figures are computed on the pulled rows and hold, while counts of unique domains are floors. Gemini ran in 16 projects and Google AI Mode and Claude in one each, so the engine findings are three-engine findings.
The August study measured 91 projects from April to August 2026 on top-30 source tables, and this study measures 138 projects across everything they ever ran, on tables up to 5,329 rows deep. None of the differences below is a trend over time. The August study found reddit.com among the top-cited sources of 83.5% of projects (76 of 91). The same top-30 cut of this dataset gives 60.9% (84 of 138), because this dataset adds many Lithuanian local-service projects whose answers cite local business sites, and the full tables give the 92.8% (128 of 138) this paper quotes.
The own-website share depends on table depth. The August study found own domains at 4.6% of 497,883 citations on top-30 tables (n=91), and the same cut of this dataset gives 6.7% of 1,220,141 citations (n=138). On the full tables the figure is 3.7% of 2,283,923, lower because deeper tables add long tail domains that belong to third parties, and this paper quotes the full-table figure with that basis attached.
The comparison-page ratio is only stable on a top-30 cut. The August study found a median citation rate of 1.70 for comparison pages against 1.06 for homepages, a 1.6x gap (n=2,730 top-cited URLs, 91 projects). The same top-30 URL cut here gives 1.47x (1.467 against 1, n=4,140 URLs, 138 projects). On the full deep tables the two medians collapse and invert (0.641 comparison against 0.667 homepage), because the tail is full of pages cited once or never and comparison pages have proportionally more of that tail. This paper makes no page-type recommendation from that ratio for that reason.
How to cite this study
Agenzy (2026). Which sources do AI engines cite? 2,283,923 citations across ChatGPT, Google AI Overviews and Perplexity. Agenzy, Vilnius. https://www.agenzy.lt/blog/which-sources-ai-engines-cite-2026
Any figure in this study may be quoted with attribution to Agenzy and a link to this page. Data window: 2025-01-01 to 2026-09-07. Instrument: Peec AI. Dataset: 138 tracked market projects, 130 businesses, 29 countries, 394,180 AI answers, 2,283,923 citations.
This URL is permanent and any correction is made on this page, in place, with a dated line in the log below.
Log: 2026-09-08, first publication.
FAQ
Which websites do AI engines cite the most?
Reddit leads, with 2.1% of the 2,283,923 citations in Agenzy's 2026 source study, and YouTube follows at 1.6%. The list is flat below that: the top 50 domains together hold 16.8% of all citations, and it takes 1,057 domains to reach half. Across 138 European business projects, no single website holds a large share.
Does ChatGPT cite Reddit more than Google AI Overviews does?
Yes. ChatGPT gave Reddit 3.0% of its citations in Agenzy's 2026 source study, Google AI Overviews gave it 1.7% and Perplexity 0.6%. Reddit also shows up in more ChatGPT projects: 77.5% of them, against 61.3% on Google AI Overviews and 46.6% on Perplexity (n=138, 137 and 133 projects). It is the most cited domain overall at 2.1%.
Do ChatGPT, Google AI Overviews and Perplexity use the same sources?
Mostly no. When the three engines answered the same questions about the same brand, a median of 2 of their top 10 cited domains were common to all three (Agenzy 2026 source study, n=109 business projects). ChatGPT and Google AI Overviews shared a median of 4, and in 5 of 119 projects that pair shared nothing.
How often do AI engines cite a brand's own website?
Brands' own websites earned 3.7% of the 2,283,923 citations in Agenzy's 2026 study, while the competitors each project named earned 16.9%. The own website was cited somewhere in 90.6% of the 138 projects, at a median rank of 10 among that project's sources, and it was missing from the 30 most cited sources in 34.8% of projects.
Does Google AI Overviews cite YouTube?
Yes, and far more than the other engines. YouTube earned 4.7% of Google AI Overviews citations in Agenzy's 2026 source study and appeared in 72.3% of the 137 projects on that engine. ChatGPT gave it 0.03%. Perplexity sits between the two, sending 0.6% of its citations to YouTube across 46.6% of 133 projects.
About this study
Agenzy is a founder-led generative engine optimization (GEO) agency in Vilnius, Lithuania. It tracks how European businesses appear in ChatGPT, Google AI Overviews, Gemini and Perplexity. The study was designed and run by Agenzy on the agency's own Peec AI tracking projects, with the aggregation script and the raw extracts kept for audit. Agenzy is an official Peec AI partner. Questions about the dataset can be sent through agenzy.lt.





