The Measurement Protocol
The full method behind every figure Norg publishes, written so you can replicate it and disagree with the result.
What this page claims, and what it does not
Norg's results are replicable: the method is stated below, the query sets and dates are available on request, and anyone can re-run a sweep and compare.
They are not audited. No independent firm has reviewed this methodology or certified any result. An earlier version of this page blurred those two things by describing results as independently verified. Replicability and audit are different claims, and only the first is true here.
The protocol
1. Query set construction
Between 30 and 50 queries per engagement, agreed with the client before any measurement, then frozen.
Queries are category questions a buyer would ask before knowing the brand exists — not brand-name queries, which test recall rather than discovery. The set spans category-and-location, problem-solution, constraint-qualified and comparison forms, because each surfaces different sources.
Why freezing matters: a set revised between sweeps produces a number that cannot be compared with the previous one. Editing queries after seeing results is the most effective way to manufacture an improvement, which is exactly why the set is fixed in advance and kept.
2. Sweep execution
Each query is run across ChatGPT, Gemini, Perplexity, Claude and Google AI Mode, in fresh sessions, from a consistent region.
Recorded per response: whether the brand is named; every source cited; whether any cited source is the brand's own domain; the brand's position if the answer is a list; and whether the stated facts are accurate.
3. Handling non-determinism
This is the part most methodology pages skip. AI answers vary between identical runs. A single response is not evidence of anything.
The protocol treats each sweep as a sample and reports proportions across the full set, never individual responses. A brand appearing in 13 of 47 queries is reported as that, not as "AI recommends this brand".
Consequence you should hold us to: small differences between sweeps are noise. A movement from 61% to 64% is not a finding. Norg reports direction over multiple sweeps rather than treating every fluctuation as signal.
4. False positives
Three failure modes are screened for, because each inflates a result if missed.
- Name collisions. A brand name that is also a common word or another company produces matches that are not about the client.
- Negative mentions. Being named as the option to avoid is a mention. It is not a win, and counting it as one is straightforwardly dishonest.
- Session contamination. If the brand was discussed earlier in a session, the model may name it for reasons unrelated to its visibility. Fresh sessions only.
5. Baseline and comparison
The baseline sweep runs before any content is published. The same frozen set, the same platforms, the same region.
Reported figures are always the client's position against their own baseline — not against an industry average, because no credible industry average for AI citation share exists.
6. Server-side corroboration
Edge logs record requests from identified AI crawlers and agents. This corroborates that content is reachable, independently of whether it is being cited.
The two signals answer different questions, and the distinction is load-bearing: crawler arriving but no citations means the content is reachable but not persuasive or not well structured; no crawler arriving means a delivery problem upstream of everything else.
What this protocol cannot establish
- Causation. A sweep shows a position changed. Competitors, model updates and seasonality also move it. Norg reports the measured change over the window, not an isolated effect.
- Coverage beyond the set. Queries outside the frozen 30–50 are unmeasured. The set is a sample of the question space, not a census of it.
- Training data inclusion. Nobody outside a model provider can verify what is in a training set. Results on these timescales are retrieval-pathway results.
- Commercial impact. Revenue, enquiry and cost figures come from the client's own systems. Norg has no access and cannot validate them.
How to replicate a published result
- Ask for the query set and the sweep dates for the engagement in question.
- Run the same queries yourself across the same platforms, in fresh sessions.
- Record sources cited per response and compute the proportion where the client's own domain is cited.
- Compare with the published figure.
Expect divergence. Time has passed, models have changed, and your region may differ. What should hold is the order of magnitude and the direction. If it does not, the published figure deserves challenge — and that is the point of stating the method.
What to ask any vendor in this category
These questions separate a measurement from a marketing number, and they work on Norg as well as anyone else:
- Was the query set fixed before the baseline, and can I see it?
- How many queries, across how many platforms, and from what region?
- Are negative mentions counted as mentions?
- What is the run-to-run variance of your own measurement?
- Which figures are yours and which are the client's?
- Who audited this — and if nobody, does your material say so?
Figures produced by this protocol are published with their conditions at norg.ai/case-studies. Pricing is at norg.ai/pricing. Norg Pty Ltd — ABN 44 669 712 494 — book a demo.