What “observed” means: how we count a mention
Every rate in Bluemoon is built from stored answers. Here are the definitions behind the count, the limits of a daily sample, and the short list of things we infer instead of observe.
When Bluemoon says a brand was “named in 31% of 547 answers”, five words are doing precise work: question, market, answer, mention and role. This note defines them, explains what one collection a day can and cannot tell you, and lists the handful of values we infer rather than observe. It is the fine print behind every screen, written so you can argue with it.
Five words we use precisely
A tracked question is a buyer question you chose to follow, written the way a buyer would type it, in the language of a market. “Best waterproof hardshell for alpine climbing” is a tracked question. A keyword is not.
A market is a country and a language together. Switzerland in German and Switzerland in French are two markets, because the engines answer them differently and cite different sites. Answers for a market are collected from inside that country.
An answer is one complete response from one engine to one tracked question in one market on one day, stored in full with the sources it displayed. It is the unit everything else is counted over.
A mention is a brand name appearing in an answer. It is detected automatically and can be reviewed by a person; when a reviewer overrules the detector, the corrected value is what counts.
A role says what the mention does. Recommended: named as a pick, with a reason. Listed: named among options with no preference. Mentioned: named in passing, as a comparison point or a qualifier. Caveated: named with a warning attached. Absent: the brand does not appear, which is recorded as an observation, not left blank.
A citation is a source the engine displayed with its answer, recorded as the URL shown and then grouped by domain. A page can be cited without any brand being mentioned, and a brand can be mentioned with no citation at all. We keep the two facts separate.
| Term | Counted as | Not counted as |
|---|---|---|
| Answer | One engine, one question, one market, one day | A conversation with follow-ups |
| Mention | The brand name, or an unambiguous product name, in the answer text | A brand appearing only inside a cited URL |
| Role | One of five values per brand per answer | A sentiment score |
| Citation | A source the engine displayed | A page the engine may have read but did not show |
One collection a day, and why that matters
Bluemoon asks each tracked question to each engine once a day per market. That gives one answer per engine per market per day, no more. It is a deliberate choice: a stable cadence makes days comparable, and it keeps the sample honest about what it is.
It also means every answer is one draw from a distribution. Ask the same engine the same question twice in a minute and you can get different brands, in a different order, with different sources. Engines are not deterministic, and they change their models and retrieval without notice. So a single day’s answer is not “what the engine says”. It is what the engine said, once, at a recorded time.
This is why the product never reports a rate from one answer, and why a change between Tuesday and Wednesday is not yet a change. The day is a data point; the finding lives in the count over days.
How many answers before we show a rate
A percentage appears only when it can carry its denominator, and only when that denominator is large enough to mean something. Below that threshold you see the count, for example “named in 4 of 11 answers”, and no percentage. Above it, the percentage and the count appear together, always.
The threshold is not magic. It is set at roughly the point where a rate stops swinging wildly with one extra answer and starts to say something about the underlying pattern. A rate at the threshold is still rough; the figure below shows how rough.
Rule of thumb, not a measurement: roughly 1 divided by the square root of the count, for a rate near 50%. Rates near 0% or 100% move less in absolute points. Use it to decide whether a difference deserves a meeting.
Thinking about uncertainty without faking precision
We do not print confidence intervals next to every rate. The arithmetic assumes independent draws, and daily answers from one engine are not quite independent: a model update moves all of them at once. A printed interval would look more exact than it is. Instead we give you the count, and a way to think.
- Read the count before the rate. 31% of 547 and 31% of 32 are different findings.
- Compare like with like: the same engines, the same market, the same period, a similar count on both sides.
- Ask what else changed. A step change on one day across every question usually means the engine changed, not you.
- Look at roles, not only presence. Moving from mentioned to recommended is the change that matters, and it needs its own count.
- Treat a single week’s movement as a hypothesis. Confirm it over the next collection window before anyone rewrites a plan.
What is observed and what is inferred
Observed means it is in a stored answer and you can open it: the question, the engine, the market, the timestamp, the full text, every displayed source, every detected mention and its role. If a rate is observed, every answer behind it is one click away.
Inferred means we computed it from something other than the answers, and it is labelled as an estimate wherever it appears. There are three of these today, and we would rather list them than let you guess.
- Intent weight. How much a question matters commercially. It is derived from public search-advertising signals and from our own judgement about the question, so it is an inference, labelled as one wherever it appears and kept out of every observed rate.
- Prompt volume. How often a question is put to the engines. No engine publishes this. Where we show a volume band it is an estimate built from public demand signals, and it is a band, not a number.
- Influence on purchases. Whether an answer changed a buying decision. We do not measure this and do not show it as a rate. Anything on this topic in Bluemoon is framed as a question for your own analytics, not a finding of ours.
The rule that follows is simple: an inferred value never sits next to an observed rate as if it were the same kind of thing. It has its own label, its inputs are listed, and a dash appears where we have nothing rather than a placeholder. The current list of what is observed and what is inferred lives on the research page and changes only when the method does.
If you find a number in the product that breaks any of these definitions, send it to us with the screen it came from. We would rather fix the count than defend it.



