How much AI answers change between runs: 30 questions, 10 runs each
We asked ChatGPT and Google AI Mode the same 30 buying questions (“best CRM for a small business”, “best password manager”, “best running shoes for flat feet”) ten times each, from the US, through the AnswerLine API. For each question we tracked six well-known brands of its category and the sites each answer cited. 596 of the 600 answers came back.
The short version: a single run tells you whether a brand can appear, and very little about how often it does.
What we found
| ChatGPT | Google AI Mode | |
|---|---|---|
| Answers collected | 297 of 300 | 299 of 300 |
| Brand and question pairs where the brand appeared at least once | 158 | 137 |
| Of those, the brand appeared in some runs and not others | 66% | 52% |
| Appeared in 20% to 80% of runs (a coin toss on any single run) | 39% | 24% |
| Cited domains shared by two runs of the same question | 30% | 59% |
Take “best password manager”. Over ten ChatGPT runs, 1Password and Bitwarden appeared every time, Dashlane eight times, Keeper twice, NordPass and LastPass never. Over ten AI Mode runs, NordPass appeared every time, Keeper nine times and Dashlane never. One run of one engine is a thin basis for telling a client where they stand.
The cited sources moved even more than the brands. Two ChatGPT answers to the same question shared under a third of their cited domains on average. A report built on one run’s sources would show a different outreach list every time you ran it.
How many runs you need
A mention rate is a share of runs, and its uncertainty shrinks with the number of runs. For the rates we observed, the 95% interval around the measured rate was on average:
| Runs per question and engine | Width of the 95% interval |
|---|---|
| 1 | about 79 points |
| 5 | about 50 to 52 points |
| 10 | about 36 to 39 points |
Ten runs are the practical floor for a number you put in front of a client, and they still leave a wide range. For a decision between two brands that sit close together, run twenty or more, and compare intervals, not single rates.
How to measure it yourself
- Repeat each prompt as independent tasks. Send the same prompt as separate async tasks in one batch; each is its own answer.
- Spread the runs out. Answers collected in the same minute share more of their conditions than answers a day apart. Split the runs over hours or days.
- Report a rate with its range. “Named in 7 of 10 answers (40% to 89%)” is a measurement. “Named” is an anecdote.
- Treat sources as a distribution. Count how many runs cite each domain and work on the ones cited most often when your brand is missing.
The measuring visibility guide shows the plain-HTTP version and the SDK helpers that do all four steps: visibility.measure runs the samples and brandCheck returns each brand’s mention rate with its 95% interval, its average rank, how many answers cite its own site and the domains cited when it is absent, with the matched text for every mention.
Method
- 30 buying-intent questions, six tracked brands per question, country US, no other targeting.
- 10 runs per question and engine, sent as independent async tasks in one batch on 2026-09-28, so this measures how much an answer varies within a few hours. Runs spread over days are likely to vary more.
- A brand counts as mentioned when one of its names appears as a whole word in the answer text, ignoring case, typographic quotes and dashes. Brands are tracked by names that are not everyday words in these answers (“Copilot Money”, not “Copilot”).
- Source overlap is the Jaccard index of the cited domains of two runs, averaged over every pair of runs of a question.
- Interval widths are Wilson 95% intervals at the observed rates.
The per-question data, every brand’s mention rate and each question’s source overlap, is in answer-variance-2026-09.json.
Questions
How many times should I run a prompt to report a mention rate?
Ten runs per prompt and engine is a practical minimum. In this study a single run left the true rate anywhere in a range about 79 percentage points wide; ten runs narrowed it to about 36 to 39 points. Spread the runs over hours or days rather than sending them in the same minute.
Do ChatGPT and Google AI Mode vary equally?
No. On ChatGPT, 66% of the brands that appeared for a question appeared in some runs and not others, and two runs of the same question shared 30% of their cited domains on average. On AI Mode the figures were 52% and 59%.
Where is the data?
Every question's mention rates and source overlap are in a JSON file linked from the article, collected on 2026-09-28 from the US.