Sample size (N)
Updated 2026-09-15
Sample size, usually written N, is the number of times each prompt is run on each engine within a measurement cycle. Because AI answers are non-deterministic, N is what converts a series of individual answers into a rate with a usable margin of error.
How N changes the number
With N = 1, a prompt is a coin flip: the brand either appeared in that single answer or not, and the resulting rate moves in whole steps. Raising N averages the variance down — the improvement is steep from one to three and gradually flattens after that — while cost rises linearly with every added sample. The practical rule is to pick the smallest N whose cycle-to-cycle noise is smaller than the change you want to detect, and to report N next to every rate.
Illustrative variance (example numbers)
A prompt whose true mention probability is 0.5 returns either 0 or 1 at N = 1. At N = 3 it returns 0, 0.33, 0.67 or 1; at N = 10 it clusters near 0.5. Nothing about the brand changed between those three measurements — only the instrument did.
Common mistakes
- Comparing two cycles measured at different N and reading the difference as a change in visibility.
- Reporting a rate without its N, which hides how much of the number is noise.
- Raising N everywhere instead of where it matters — cost scales with every prompt on every engine.
Frequently asked questions
- What is a reasonable N for weekly tracking?
- Three per prompt per engine is a common working floor and keeps the cost predictable; going higher is worth it on the handful of prompts whose movement drives decisions.
- Does a higher N make the number more "true"?
- It makes it more stable, which is different. The rate still describes what these engines did with this prompt set during this cycle — a proxy measurement, honestly bounded.
