AI answers resample — why one check is an anecdote
Run “best project management tools” through ChatGPT twice and you may get different recommendations, different sources, different ordering. This isn’t a bug to work around — it’s the defining property of the medium, and it changes how measurement has to work.
Why answers vary
- Sampling: generation is stochastic — the model draws from a distribution, not a table
- Retrieval drift: the fan-out searches return whatever ranks today; the cited pool shifts with the index
- Personalization and geo: account history, market and locale all feed the answer
- Model updates: silent version changes can restyle answers overnight
What variance means for measurement
A single run tells you what can be said, not what is said. The honest metrics are rates over samples:
- Mention rate over N runs per prompt, not a boolean
- Citation share across the prompt set, not per run
- Trends computed weekly from repeated samples, not day-over-day deltas that are mostly noise
Practical guidance: 3 runs per prompt-market-engine cell for directional work, 30+ before publishing a number. The cadence post covers scheduling; share of voice has the formulas.
How to report it
Show variance, don’t hide it: a brand with 60% mention rate and wide error bars is “sometimes recommended” — that’s the finding, not a data quality problem. Stakeholders trust a number more when they can see its range.
When variance itself is the signal
A prompt where answers flip between you and a competitor run to run is a contested intent — winnable with better-cited content. A prompt with a locked answer is either owned or lost. Variance maps tell you where effort can still move the needle — collect the samples with a scheduled batch and let the rates settle.