Key Takeaways
- There is no single "AI search result." ChatGPT, Perplexity, Gemini, Claude and Google's AI features behave like separate products.
- In a September 2026 study by Prefer, an AI-visibility software vendor, 72.7% of the sites that AI engines cited appeared in only one engine. Only 2.2% appeared in all four. The study has limits, covered below.
- The same question asked twice returned very different sources on some engines. Repeat-run overlap in the study: Perplexity 96%, Claude 68%, ChatGPT 37%, Gemini 35%.
- Not every answer searches the web. In the study, the 99 answers that skipped search cited nothing. Record whether each run searched.
- A fair check uses a fixed list of buyer questions, at least three runs per question per platform, a separate record for each platform, and a repeat. We recommend monthly, and the trend matters more than any single answer.
- As of September 2026, Google says SEO fundamentals still apply to its AI features and points owners to a Search Console report. The Search Console Generative AI performance report shows how your content performs in those features.
What goes wrong when you check ChatGPT once?
A common test goes like this. An owner opens ChatGPT, types "best patio builder in my city," and reads the answer. If the business is named, they relax. If not, they worry. Both reactions rest on one data point, and it is weaker than it looks.
A single AI answer often gets treated as proof in both directions. A screenshot goes into a report. A slide says "visible in ChatGPT." Nobody asks how many times it was run, on which platforms, or whether the tool searched the web at all.
The one-check method has three problems. Published data backs each one.
What does the research say about AI platforms and citations?
The best public dataset we have found is from Prefer. Prefer sells software that tracks AI citations, so it has a stake in the topic. It also published its method and raw data, which is why we are using it.
What the study did. Prefer, an AI-visibility software vendor, published a study on September 19, 2026 [1] that put 80 questions about AI search to four engines, three times each: ChatGPT, Perplexity, Gemini and Claude. That is 960 answers. Data was collected on September 13, 2026, through the DataForSEO LLM Responses API, with web search requested on every call. The answers cited 1,329 different sites. You can read the full write-up on Prefer's study page [1].
The engines mostly cite different sites. In Prefer's data [1], of the 1,329 sites cited, 966 (72.7%) were cited by one engine only. Only 29 (2.2%) were cited by all four. On 69 of the 80 questions, no single site was cited by all four engines.
The big sites are the exception. Prefer also found [1] that 23 of the 25 most-cited sites were cited by three or four engines. So the engines tend to agree on the biggest names and disagree on most of what sits below them. A headline that says "73% of citations are unique" is true and incomplete. Half the story is the shared top of the list.
Engines differ in how much they cite. Prefer measured average sources per answer [1]: Perplexity 19.48, Gemini 8.64, Claude 4.58, ChatGPT 3.05.
The same engine does not repeat itself equally. Prefer [1] compared two runs of the same question. They shared this share of cited sites on average: Perplexity 96%, Claude 68%, ChatGPT 37% and Gemini 35%. Prefer divided the sites cited in both runs by all sites cited across both, averaged over run pairs, and excluded pairs where a run cited nothing. Counting those pairs as zero gives Perplexity 96%, Claude 59%, ChatGPT 34% and Gemini 31%. The direction holds.
Engines do not always search. In Prefer's study [1], web search was requested every time. The engines still chose to use it on 100% of answers for Perplexity, 90% for ChatGPT, 88.8% for Gemini and 80% for Claude. All 99 answers that skipped search cited no sources. As of September 2026, OpenAI's help page on searching the web with ChatGPT [2] says ChatGPT "may search the web automatically when your question would benefit from current information."
Each engine leans on different places. In Prefer's data [1], Reddit appeared in 207 of 960 answers and never in Claude's. All 113 answers citing YouTube came from Gemini.
What are the limits of this study?
Prefer states these limits itself [1]. Treat the numbers as one well-documented snapshot, not a law.
- It has one collection date.
- The 80 questions were about AI search itself, so other topics can behave differently.
- It counted what the engines cited, not whether the answers were correct.
- It used API answers. What a person sees in the consumer apps can differ.
- It measures variation within minutes, not change over weeks.
- Patterns move. Prefer's earlier July 2026 study [3] used 47 buyer questions and found that 5 of 713 domains were cited by all four engines, a different mix from the September study.
Do not quote these figures as your market's numbers. They show one thing: why a single check cannot be trusted.
How do you measure AI search visibility fairly?
The method is dull on purpose. It has five steps.
1. Write a fixed list of buyer questions
Pick 10 to 15 questions a real buyer asks before hiring a business like yours. Mix the stages:
- Problem questions, such as "why are my Google Ads leads so expensive?"
- Comparison questions, such as "Google Ads or SEO for a service business?"
- Hiring questions, such as "how do I choose a marketing agency?"
- Local questions, if you serve a defined area
Keep the wording identical between rounds, or you cannot compare to last month. Put the questions closest to revenue first. "Who should I hire to build a patio in my city?" is closer to a phone call than "how much does a patio cost?" The post on marketing metrics that predict revenue explains why that matters.
2. Run every question at least three times on every platform
Use a fresh chat each time so earlier answers do not shape later ones. Note the date, the platform and, if the tool shows it, the model. Prefer also recommends at least three runs per question on each engine [1]. The goal is to see whether an answer is stable.
3. Record whether the tool searched the web
Write yes or no for every run. A run with no sources does not always mean you were dropped. The tool may have answered from memory, and mixing those cases misleads.
4. Report each platform on its own line
Do not blend everything into one "AI visibility score." A blended number hides which platform, for which question, cites whom. A simple layout works:
| Question | Platform | Run | Searched the web? | Sources named | Were you named or cited? | Notes |
|---|---|---|---|---|---|---|
| (your fixed question) | ChatGPT | 1 | Yes / No | (list) | Named / Cited / Neither | |
| (your fixed question) | ChatGPT | 2 | ||||
| (your fixed question) | Perplexity | 1 |
Keep three things separate. Industry tools define a mention (your name appears in the answer) and a citation (your page is listed as a source) as different things, and they often appear separately [8][9]. In a Semrush study of 3,981 domain appearances, 61.7% were cited without being named in the answer, 13.2% were both, and 25.1% were named without a citation [10]. We also track whether the citation is a clickable link, and you can note that in the last column.
5. Repeat it, we recommend monthly, and read the trend
We recommend a monthly check, run several times per prompt on each engine, because AI answers change from run to run and week to week. Read the trend, not a single answer. In SISTRIX's study of 82,619 prompts over 17 weeks, Google replaced 56% of the sources in AI answers every week and ChatGPT as much as 74% [6]. In Profound's comparison of June and July 2025, 40% to 59% of the cited domains had changed a month later, depending on the engine [7]. Both are software vendors. No study we found tests whether monthly is enough, so treat monthly as our recommendation, not a finding.
The first round is a baseline. The second is where you learn something. Look for:
- Questions where you appear on most runs on most platforms. That is a strong position.
- Questions where you appear on one platform only. Find out what that platform reads.
- Questions where the same competitor or directory keeps appearing. That is the source set you are up against.
- Changes that hold for several months. Single-run swings are mostly noise.
Report rates with their denominators, for example "named in 4 of 12 runs," not a score. Skip targets like "appear in half of all answers." With this much variation, a target invites you to chase noise.
What does Google say you should do?
As of September 2026, Google's guide to optimizing for generative AI features [4] (last updated July 10, 2026) says SEO best practices continue to matter because these features are rooted in its core Search ranking and quality systems. It treats "AEO" and "GEO" as SEO. It says you can ignore tactics like chunking content, creating unnecessary AI text files such as llms.txt, or pursuing inauthentic mentions. It says structured data isn't required for generative AI search and that no special schema.org markup is needed. And it points owners to the Generative AI performance report in Search Console to see how their content performs in those features.
For Google's own surfaces, your Search Console data beats any screenshot.
It also means you should be careful with anyone selling a separate trick for AI answers. Google's changelog [5] says the FAQ rich result stopped appearing in Google Search on May 7, 2026. Google's guide [4] gives no role to FAQ markup for generative AI features. On our pages, FAQ sections are there for readers. If a reader finds the answer fast, the page is doing its job.
How do AI answers fit with the leads you actually need?
AI answers are one surface to watch. The Map Pack, Local Service Ads, Google Ads, referrals and reviews are others. The post on what AI Overviews mean for service-business leads covers how this plays out.
Ask of any AI visibility number what you ask of every marketing number: does it connect to revenue? A mention in an AI answer is not a lead. A lead is a call, a form or a booked job. Tie each round of the check to what happens downstream, as in the post on marketing attribution for service businesses.
The foundation has not changed. Clear pages and real expertise still do the work. See E-E-A-T and trust for service businesses and content marketing that compounds.
What should you do with the results?
Read them for patterns, not scores. If the tools keep citing a directory or a competitor's guide for a question you should own, write a clearer page on that question. Fix the obvious gaps, re-run the same list next month and write down what changed.
A note on vendors. Plenty of tools will run prompts for you and show a dashboard. Prefer, the source of the data above, is one of them. Ask any AI-visibility vendor three things: does it repeat each prompt, does it test the web interface or the API, and does it show whether the engine searched the web. Vendors differ on all three, and studies show the API and the chat window can give different answers [11][12]. A tool that cannot answer all three is giving you a screenshot.
Where should you start?
If you want to see where your current marketing stands first, the Revenue System Scorecard gives you an instant score with no sales call. If you are ready to work together, apply here. We reply within 1-2 business days. For what each step includes, see pricing.
