Why do AI answers change every time I ask?
Short answer: Models are probabilistic and retrieval varies run to run, so one prompt can give different lists. Measure with repeated samples, not a single answer.
Because the models are probabilistic and the retrieval step is not fixed. Ask ChatGPT "best wedding photographer in Cardiff" three times and you may get three overlapping but different lists. That is normal, and it is the single biggest reason manual spot checks mislead.
Three mechanisms cause the variation:
- Sampling. A language model chooses each word from a probability distribution. Even at low "temperature" settings, small differences early in the answer cascade into a different set of names at the end.
- Retrieval. Engines that search the web before answering (ChatGPT with search, Gemini, Perplexity, AI Overviews) run slightly different queries and get slightly different pages each time. If your page is on the edge of what gets retrieved, you appear in some runs and not others.
- Location and session. Consumer apps add personalisation, memory and location. Two people in different towns, or the same person on different days, see different answers.
The consequence for measurement is that "we were mentioned" or "we were not" is the wrong unit. The right unit is the proportion of runs in which you are mentioned. Being named in two of three samples this week and three of three next week is progress; being named once in a single manual check tells you nothing.
Seoptist runs each prompt three times per engine every week (five on Enterprise), averages the results, and shows a confidence band around the visibility line so you can tell a real shift from noise. Repeated runs within 24 hours for the same prompt, engine and location are cached rather than re-asked, which keeps cost down without hiding the variation that matters.
If you are tracking manually, ask at least three times, in a fresh session, and record all three.