LLM measurement runs
AI answers are non-deterministic. Ask ChatGPT the same prompt twice and you can get two different brand lists, in a different order, with different citations. That's a feature for users β and a problem for measurement.
To produce a number you can trust week over week, GeoBubbles runs each prompt multiple times per measurement cycle and averages the result.
How a cycle works
When a measurement cycle runs for your site:
- SEO data (positions, search volume, SERP features) is collected once per cycle β these are deterministic.
- Each chat assistant is queried 3 times per active prompt during the cycle. We use the same prompt text every time, but each call goes out as a fresh, isolated session.
- The 3 runs are averaged into a single per-prompt score for that cycle, and that average is what feeds the dashboard, charts, and reports.
A side-effect: one weird outlier (e.g. one run where the AI hallucinated and skipped every brand) gets smoothed out instead of dragging your score down.
Why 3 runs and not 1 or 10
We settled on 3 after benchmarking variance across thousands of prompts:
- 1 run β too noisy. A single bad sample can flip a brand from "always cited" to "missing" day over day.
- 3 runs β variance drops sharply. Trends become readable; week-over-week changes reflect real movement, not sampling noise, while keeping per-cycle cost reasonable.
- 5+ runs β incremental precision gain, materially higher cost. Not worth it for most use cases.
Trial sites
Free-trial sites run 1 query per prompt per cycle instead of 3. This keeps the trial generous without burning the full per-run cost. The trade-off is more visible run-to-run variance β once you're on a paid plan, the full 3-run averaging kicks in and your scores stabilise.
Reading the dashboard
Most dashboard metrics show you the averaged value per cycle. When you drill into a single prompt's detail view, you can see each of the underlying runs (which brands appeared, in what order, with what citations) β useful for diagnosing why a metric moved.