AI search answers change constantly. Run the same prompt again tomorrow and you may see different brands, different sources, or a different answer altogether.

For marketers, that creates a basic measurement problem: which AI visibility signals are stable enough to track, and which ones should you expect to move?

In this study, we define stability as how consistently the same signal appears when the same prompt is run again over time.

We focused on two signals that matter most for AI visibility:

  • Brand mentions: which brands appear in the answer.
  • Citations: which sources the AI engine uses.

We tracked 520 prompts across three industries, running them daily across seven AI answer engines for 40 to 91 days, and collected 252,407 answers.

The clearest pattern was that brand mentions were generally more stable than the sources behind them. Citations changed much more from one day to the next, and the amount of change varied widely by engine.

That leads to our main recommendation: we would use brand coverage as the top-line KPI for AI visibility, and citations as the diagnostic layer underneath it.

Brand coverage is the share of tracked prompts in which your brand is mentioned. It tells you whether your brand is showing up. Citations help explain what information and sources may be supporting that visibility.

Key findings

  • Brand mentions are generally more stable than citations. Across the six comparable engines, only 12% to 19% of brands appeared on a single day during the 30-day window and were never mentioned again. For cited sources, that share reached as high as 67%, meaning that on some engines, around two thirds of sources appeared on just one day during the same period.
  • More prompts make Brand Coverage more consistent. In our 320-prompt dataset, moving from 10 to 100 prompts reduced variation in the reported Brand Coverage by about 73%.
  • Citation sources turn over quickly, and the pattern depends heavily on the engine. In our AI search monitoring dataset, ChatGPT and Google AI Mode retained only around a quarter of the sources they had cited the day before, while Perplexity retained roughly three quarters.
  • A stable brand signal can hide a much more dynamic source layer. Your brand can continue appearing at a similar rate while the pages and domains supporting those answers change underneath it. That is why we see citations as a diagnostic signal rather than the headline KPI.
  • Nearly half of all URLs were cited only once. Across all seven engines, 48% of the 45,919 URLs appeared just once during the 30-day window, while the top 1% accounted for 44% of all citations.
  • Each engine behaves differently. The differences are not limited to citations. Engines also showed recognizable answer patterns in length, table usage, and whether they ended answers with a question. Those patterns were broadly consistent across the categories we measured.

Brand coverage is the KPI we would report first

Brand mentions and citations do not move at the same rate.

Across five of the six engines we could compare fairly, a brand mentioned today was more likely to appear again tomorrow than a source cited today was to be cited again. Claude was the exception, where the two were almost equally stable. Perplexity was excluded from brand-based comparisons because 82% of its answers did not contain a brand mention.

The difference becomes even clearer when we look at one-off appearances.

For each prompt and engine, we looked at all the brands and sources that appeared during a 30-day window and measured how many appeared on only one of those 30 days.

The source layer was also much larger. For a given prompt, engines cited around 95 different sources over the month on average, compared with about 10 brands mentioned. That wider source layer also had a much longer tail of one-off appearances.

Depending on the engine, only 12% to 19% of the brands mentioned for a prompt appeared on just one day during the 30-day window. For cited sources, that ranged from 26% to 67%. On Google AI Mode, for example, about two thirds of the sources cited for a prompt appeared on only one day, compared with 17% of the brands mentioned.

Bar chart comparing one-day-only cited sources and brand mentions across six AI engines over 30 days, with cited sources showing much higher turnover.

That distinction matters for reporting.

If the question is simply “Are we showing up in AI answers?”, we would start with brand coverage. It gives a more stable top-line view of visibility and is easier to interpret from one reporting period to the next.

Citations answer a different question: “What is supporting that visibility?”

They tell you which pages and publishers are feeding the answer, but that layer changes much faster. A citation disappearing tomorrow does not necessarily mean your overall AI visibility has deteriorated. It may simply reflect normal source turnover in the engine.

So we would not treat brand coverage and citations as two interchangeable measures of AI visibility:

Use brand coverage to track whether you are visible. Use citations to investigate how that visibility is being built.

How many prompts should you track?

In our test, moving from 10 to 100 prompts reduced the variation in Brand Coverage by about 73%. The biggest gains came from adding prompts early, and the improvement became smaller as the prompt set grew.

With only 10 prompts, 90% of the Brand Coverage results fell between 39% and 71%. With 50 prompts, that range narrowed to 49%–62%. At 100 prompts, it narrowed further to 51%–60%.

Line chart showing that Brand Coverage variation decreases as the number of tracked prompts increases, from 10 to 200 prompts.

We tested this by repeatedly measuring Brand Coverage on different random subsets of our 320 observed prompts, while keeping the brand, engines and 30-day period fixed.

This does not mean that 100 prompts is the right number for every brand. But in our dataset, small prompt sets produced less reliable Brand Coverage measurements, while broader prompt sets gave a more consistent result.

Brand position was relatively stable too

Brand coverage was not the only brand-level metric that showed relatively little movement.

We also tracked OtterlyAI’s average brand position in answers over the same period. Across the six engines where there was enough brand mention data to compare, the average position moved by only around 0.09 to 0.15 places per day, with no clear upward or downward trend over the month.

EngineAverage placeBest to worst dayMoves per day
Google AI Overviews2.352.17 to 2.580.09
Copilot2.962.65 to 3.230.10
ChatGPT3.312.83 to 3.850.10
Gemini3.072.83 to 3.610.13
Claude3.162.89 to 3.520.13
Google AI Mode3.202.81 to 3.720.15

This suggests that brand position can work as a useful supporting metric alongside brand coverage. Brand coverage tells you whether your brand is appearing. Position adds context on how prominently it appears when mentioned.

Citation churn is normal, but it varies a lot by engine

Citation data is much more volatile than brand coverage.

In our AI search monitoring dataset, more than half of the sources cited by ChatGPT, Gemini and Google AI Overviews appeared on only one day during the 30-day window. On Google AI Mode, that rose to 66.5%. Claude and Perplexity were much less volatile, at around a quarter.

Bar chart showing the share of cited sources that appeared on only one day across seven AI engines, ranging from 26.1% for Claude to 66.5% for Google AI Mode.

We see the same pattern when we compare one day directly with the next. ChatGPT and Google AI Mode reused only about 26% of the same sources, while Perplexity reused roughly 75%.

EngineOtterlyAIGovernmentRetail
Perplexity74.9%73.3%65.2%
Claude62.3%64.1%75.6%
Copilot46.2%50.9%25.9%
Google AI Overviews36.1%22.1%29.2%
Gemini28.4%21.0%22.3%
Google AI Mode26.4%20.9%22.5%
ChatGPT25.8%27.4%28.9%

That is a wide range, and it changes how citation movements should be interpreted.

If one of your pages is cited today and disappears tomorrow, that is not automatically evidence that you lost visibility because of something you changed. On some engines, source turnover is simply part of normal behaviour.

The reverse is also true. A new citation is not necessarily proof that an optimization worked.

This is why we would treat citation data as a diagnostic signal rather than a standalone performance verdict. The useful question is not only whether a citation appeared or disappeared, but whether the movement is sustained and whether it is accompanied by a broader change in brand visibility.

Nearly half of cited URLs appeared only once

The same pattern becomes even clearer when we look at how often individual URLs were cited.

Across all seven engines, we found 45,919 distinct URLs during the 30-day window. Almost half of them, 48%, were cited only once.

At the other end of the distribution, citation activity was highly concentrated. The top 1% of URLs accounted for 44% of all citations, while the top 10% accounted for 81%.

Log-scale chart of citation frequency across URLs, showing that 48% were cited only once while the top 1% accounted for 44% of all citations.

This shows the long tail behind citation data: AI engines draw from a very large pool of URLs, but most of those URLs appear only occasionally, while a relatively small group is cited repeatedly.

For marketers, that is another reason not to overinterpret a single citation. Appearing once is common. The more useful signal is whether a page becomes part of the smaller group of sources that engines return to repeatedly.

These numbers are not universal benchmarks

There is another reason to be careful with citation benchmarks: the same engine did not always behave the same way across the three categories we measured.

For example, ChatGPT was remarkably similar across AI search monitoring, public sector contracting and retail. Other engines moved much more between categories.

So we would not take a number such as “ChatGPT retains 26% of its sources” and treat it as a fixed property of ChatGPT.

It describes what we observed in this dataset.

Measure your own baseline before deciding whether a citation change is unusual.

Each engine structures answers differently

The engines did not just differ in which brands and sources they used. They also had very different ways of presenting an answer.

Claude and Copilot used tables in 87% of answers. ChatGPT was also table-heavy, at 76%. At the other end, Google AI Overviews and Perplexity used tables in only 1% of answers.

Comparison of AI answer formats across seven engines, showing use of tables, answers ending with a question, and average answer length in characters.

The difference is just as striking at the end of the response. Copilot finished 83% of its answers by asking the reader a question. ChatGPT did so in just 0.3%.

Answer length varied too. Google AI Mode produced the longest answers on average, at around 7,300 characters, while Perplexity averaged about 2,500.

Comparison of the same Google AI Overviews prompt on consecutive days, showing answer length dropping from 10,821 to 4,401 characters and citations from 13 to 5.

Why does this matter for GEO?

Because “AI search” is not one uniform output format. The same content may be surfaced differently depending on the engine: inside a table, as short prose, in a much longer answer, or alongside a follow-up question.

That does not mean brands should create a separate content strategy for every engine. But it does mean that a single benchmark for “what an AI answer looks like” is misleading.

There is no single “AI answer format.” Each engine should be evaluated on its own terms.

And even these engine-level patterns are not guarantees. Individual answers still varied. On Gemini, for example, many prompts produced a table on some days and no table on others during the same month.

What this means for marketers

The data points to three practical recommendations for AI visibility reporting.

1. Use brand coverage as your top-line KPI.

If your main question is whether your brand is visible in AI answers, start with brand coverage.

Brand mentions were generally more stable than individual citations across the engines we measured. That makes brand coverage easier to interpret as a reporting metric over time.

This does not mean citations are less important. They simply answer a different question.

2. Use citations to diagnose what is happening underneath.

Citation data shows which pages and publishers are being cited in AI answers, but that layer changes much faster.

A page appearing or disappearing from citations can be meaningful, especially if the change persists. But a single citation gain or loss should not be treated as proof that an optimization worked or failed.

Look for repeated patterns rather than isolated citation changes.

3. Build your own baseline for each engine and category.

The engines behaved very differently from one another, and the same engine did not always produce the same stability level across categories.

So the numbers in this study are useful as context, not as universal benchmarks.

The same movement may be unusual in one setup and completely normal in another.

The goal is not to eliminate volatility. It is to learn what normal volatility looks like in your own data, so you can recognize when something genuinely changes.

That is also why we would avoid judging AI visibility from a single answer or screenshot. One response is an observation. A pattern across your tracked prompts is a signal.

How we measured this

We tracked the same prompts every day across seven AI answer engines: ChatGPT, Perplexity, Gemini, Google AI Mode, Google AI Overviews, Microsoft Copilot and Claude.

The study covered three categories:

  • AI search monitoring: 320 prompts tracked for up to 91 days
  • Public sector contracting: 100 prompts tracked for 40 days
  • Retail and e-commerce: 100 prompts tracked for 40 days

Together, that produced 252,407 answers. All runs were collected in the United States, with one run per prompt, per engine, per day.

Most results in this article are direct measurements from the stored answers. For the prompt-count analysis, we also used repeated random subsets of our observed 320 prompts to measure how much the reported result changed with different prompt sample sizes.

We counted which brands were mentioned, which sources were cited, how much those sets overlapped from one day to the next, and how often a brand or source appeared on only one day during a 30-day window. We also measured answer characteristics such as table usage, answer length and whether the response ended with a question.

A few limitations

These figures describe what we observed in these three categories during this measurement period. They should not be treated as fixed stability rates for each engine.

The three datasets also do not have the same history: the AI search monitoring report has 91 days of data, while the two industry reports have 40 days. For comparisons across categories, we therefore use the common 30-day window.

Finally, each prompt was run once per day. That means this study measures day-to-day variation, not how much the same engine might vary if the identical prompt were run repeatedly within the same hour.

Conclusion

AI visibility is not a single signal. Brand mentions and citations behave differently, and the amount of variation also changes substantially from one engine to another.

That is why we would put brand coverage at the top of the reporting layer and use citations to understand what is happening underneath it. The goal is not to expect every answer to stay the same. It is to understand what normal variation looks like, so meaningful changes are easier to recognize.

One AI answer is an observation. A pattern across your data is a signal.

Running a GEO experiment? Share it with us

We are continuously testing what actually moves AI Search visibility at OtterlyAI. Our public GEO Experimentation Tracker brings together experiments from our team and the wider GEO community.

If you have an experiment, result, or hypothesis worth testing, email thaylise.nakamoto@otterly.ai. If it is a good fit, we may include it in the tracker and credit you for the contribution.

👉 Want to measure your AI Visibility? Start with OtterlyAI and our GEO Guide.

👉 Learn how to connect your Claude or ChatGPT to OtterlyAI’s MCP