AnswerLine Start free

, GEO · Analytics · Research

How much AI answers change between runs: 30 questions, 10 runs each

We asked ChatGPT and Google AI Mode the same 30 buying questions (“best CRM for a small business”, “best password manager”, “best running shoes for flat feet”) ten times each, from the US, through the AnswerLine API. For each question we tracked six well-known brands of its category and the sites each answer cited. 596 of the 600 answers came back.

The short version: a single run tells you whether a brand can appear, and very little about how often it does.

What we found

ChatGPT Google AI Mode
Answers collected 297 of 300 299 of 300
Brand and question pairs where the brand appeared at least once 158 137
Of those, the brand appeared in some runs and not others 66% 52%
Appeared in 20% to 80% of runs (a coin toss on any single run) 39% 24%
Cited domains shared by two runs of the same question 30% 59%

Take “best password manager”. Over ten ChatGPT runs, 1Password and Bitwarden appeared every time, Dashlane eight times, Keeper twice, NordPass and LastPass never. Over ten AI Mode runs, NordPass appeared every time, Keeper nine times and Dashlane never. One run of one engine is a thin basis for telling a client where they stand.

The cited sources moved even more than the brands. Two ChatGPT answers to the same question shared under a third of their cited domains on average. A report built on one run’s sources would show a different outreach list every time you ran it.

How many runs you need

A mention rate is a share of runs, and its uncertainty shrinks with the number of runs. For the rates we observed, the 95% interval around the measured rate was on average:

Runs per question and engine Width of the 95% interval
1 about 79 points
5 about 50 to 52 points
10 about 36 to 39 points

Ten runs are the practical floor for a number you put in front of a client, and they still leave a wide range. For a decision between two brands that sit close together, run twenty or more, and compare intervals, not single rates.

How to measure it yourself

  1. Repeat each prompt as independent tasks. Send the same prompt as separate async tasks in one batch; each is its own answer.
  2. Spread the runs out. Answers collected in the same minute share more of their conditions than answers a day apart. Split the runs over hours or days.
  3. Report a rate with its range. “Named in 7 of 10 answers (40% to 89%)” is a measurement. “Named” is an anecdote.
  4. Treat sources as a distribution. Count how many runs cite each domain and work on the ones cited most often when your brand is missing.

The measuring visibility guide shows the plain-HTTP version and the SDK helpers that do all four steps: visibility.measure runs the samples and brandCheck returns each brand’s mention rate with its 95% interval, its average rank, how many answers cite its own site and the domains cited when it is absent, with the matched text for every mention.

Method

The per-question data, every brand’s mention rate and each question’s source overlap, is in answer-variance-2026-09.json.

Questions

How many times should I run a prompt to report a mention rate?

Ten runs per prompt and engine is a practical minimum. In this study a single run left the true rate anywhere in a range about 79 percentage points wide; ten runs narrowed it to about 36 to 39 points. Spread the runs over hours or days rather than sending them in the same minute.

Do ChatGPT and Google AI Mode vary equally?

No. On ChatGPT, 66% of the brands that appeared for a question appeared in some runs and not others, and two runs of the same question shared 30% of their cited domains on average. On AI Mode the figures were 52% and 59%.

Where is the data?

Every question's mention rates and source overlap are in a JSON file linked from the article, collected on 2026-09-28 from the US.

Try it on your own prompts

500 free credits, once — no card. One POST returns the answer, sources and citations as JSON.

Keep reading