Two runs of the same prompt can diverge because one is anonymous while the other inherits account history, location, or a selected search/grounding mode. If a dashboard blends those observations, the average looks precise but is not reproducible.
A practical split:
1. Controlled baseline fixed prompt version, market/language, engine and mode; a fresh session; timestamp; and the cited URLs.
2. Personalized diagnostic explicitly tagged as a returning session or profile, with only scenario-level interpretation.