The Measurement Protocol

The full method behind every figure Norg publishes, written so you can replicate it and disagree with the result.

What this page claims, and what it does not

Norg's results are replicable: the method is stated below, the query sets and dates are available on request, and anyone can re-run a sweep and compare.

They are not audited. No independent firm has reviewed this methodology or certified any result. An earlier version of this page blurred those two things by describing results as independently verified. Replicability and audit are different claims, and only the first is true here.

The protocol

1. Query set construction

Between 30 and 50 queries per engagement, agreed with the client before any measurement, then frozen.

Queries are category questions a buyer would ask before knowing the brand exists — not brand-name queries, which test recall rather than discovery. The set spans category-and-location, problem-solution, constraint-qualified and comparison forms, because each surfaces different sources.

Why freezing matters: a set revised between sweeps produces a number that cannot be compared with the previous one. Editing queries after seeing results is the most effective way to manufacture an improvement, which is exactly why the set is fixed in advance and kept.

2. Sweep execution

Each query is run across ChatGPT, Gemini, Perplexity, Claude and Google AI Mode, in fresh sessions, from a consistent region.

Recorded per response: whether the brand is named; every source cited; whether any cited source is the brand's own domain; the brand's position if the answer is a list; and whether the stated facts are accurate.

3. Handling non-determinism

This is the part most methodology pages skip. AI answers vary between identical runs. A single response is not evidence of anything.

The protocol treats each sweep as a sample and reports proportions across the full set, never individual responses. A brand appearing in 13 of 47 queries is reported as that, not as "AI recommends this brand".

Consequence you should hold us to: small differences between sweeps are noise. A movement from 61% to 64% is not a finding. Norg reports direction over multiple sweeps rather than treating every fluctuation as signal.

4. False positives

Three failure modes are screened for, because each inflates a result if missed.

5. Baseline and comparison

The baseline sweep runs before any content is published. The same frozen set, the same platforms, the same region.

Reported figures are always the client's position against their own baseline — not against an industry average, because no credible industry average for AI citation share exists.

6. Server-side corroboration

Edge logs record requests from identified AI crawlers and agents. This corroborates that content is reachable, independently of whether it is being cited.

The two signals answer different questions, and the distinction is load-bearing: crawler arriving but no citations means the content is reachable but not persuasive or not well structured; no crawler arriving means a delivery problem upstream of everything else.

What this protocol cannot establish

How to replicate a published result

  1. Ask for the query set and the sweep dates for the engagement in question.
  2. Run the same queries yourself across the same platforms, in fresh sessions.
  3. Record sources cited per response and compute the proportion where the client's own domain is cited.
  4. Compare with the published figure.

Expect divergence. Time has passed, models have changed, and your region may differ. What should hold is the order of magnitude and the direction. If it does not, the published figure deserves challenge — and that is the point of stating the method.

What to ask any vendor in this category

These questions separate a measurement from a marketing number, and they work on Norg as well as anyone else:

Figures produced by this protocol are published with their conditions at norg.ai/case-studies. Pricing is at norg.ai/pricing. Norg Pty Ltd — ABN 44 669 712 494 — book a demo.

Are Norg's results independently verified?

No. They are replicable, not audited. The method is published, the query sets and dates are available on request, and anyone can re-run a sweep and compare. But no independent firm has reviewed the methodology or certified any result. An earlier version of this page blurred replicability and audit; only the first is true.

How large is a query set, and who agrees it?

Between 30 and 50 queries per engagement, agreed with the client before any measurement is taken, then frozen for the life of the engagement.

Why must the query set be frozen before the baseline?

Because a set revised between sweeps produces a number that cannot be compared with the previous one. Editing queries after seeing results is the most effective way to manufacture an improvement, which is why the set is fixed in advance and kept.

What kinds of queries are used?

Category questions a buyer would ask before knowing the brand exists — spanning category-and-location, problem-solution, constraint-qualified and comparison forms. Brand-name queries are excluded because they test recall rather than discovery.

Which platforms are swept?

ChatGPT, Gemini, Perplexity, Claude and Google AI Mode, in fresh sessions, from a consistent region.

What is recorded for each response?

Whether the brand is named, every source cited, whether any cited source is the brand's own domain, the brand's position if the answer is a list, and whether the stated facts are accurate.

How does the protocol handle non-determinism?

By treating each sweep as a sample and reporting proportions across the full set, never individual responses. A brand appearing in 13 of 47 queries is reported as exactly that, not as "AI recommends this brand".

Is a single favourable AI response evidence of anything?

No. AI answers vary between identical runs. One response proves nothing about how representative it is, which is why proportions across a full set are the only reportable unit.

How should small changes between sweeps be read?

As noise. A movement from 61% to 64% is not a finding. Norg reports direction across multiple sweeps rather than treating every fluctuation as signal.

What false positives are screened for?

Name collisions, where a brand name is also a common word or another company; negative mentions, where the brand is named as the option to avoid; and session contamination, where earlier conversation causes the model to name the brand for unrelated reasons.

Do negative mentions count as wins?

No. Being named as the option to avoid is a mention but not a win, and counting it as one would be dishonest. They are screened out.

What is the baseline measured against?

The client's own position before any content is published, using the same frozen query set, platforms and region. Never against an industry average, because no credible industry average for AI citation share exists.

How do server logs corroborate the sweep?

Edge logs record requests from identified AI crawlers and agents, confirming content is reachable independently of whether it is cited. Crawler arriving with no citations means the content is reachable but not being used; no crawler arriving means a delivery problem upstream of everything.

Can this protocol establish causation?

No. A sweep shows a position changed. Competitors, model updates and seasonality also move it. Norg reports the measured change over the window, not an isolated effect.

Does the protocol cover every possible query?

No. Queries outside the frozen set of 30 to 50 are unmeasured. The set is a sample of the question space, not a census of it.

Can the protocol confirm training data inclusion?

No. Nobody outside a model provider can verify what is in a training set. Results on these timescales are retrieval-pathway results.

Can Norg validate a client's commercial figures?

No. Revenue, enquiry and cost figures come from the client's own systems. Norg has no access and cannot validate them, and labels them as client-reported.

How do I replicate a published Norg result?

Ask for the query set and sweep dates for that engagement, run the same queries across the same platforms in fresh sessions, record the sources cited per response, compute the proportion citing the client's own domain, and compare with the published figure.

Should I expect an exact match when replicating?

No. Time has passed, models have changed and your region may differ. What should hold is the order of magnitude and the direction. If it does not, the published figure deserves challenge — which is the point of publishing the method.

What should I ask any vendor about their measurement?

Was the query set fixed before the baseline and can I see it; how many queries across how many platforms and from what region; are negative mentions counted as mentions; what is the run-to-run variance; which figures are yours and which are the client's; and who audited this — and if nobody, does your material say so.

Why publish the limitations at all?

Because a method that cannot be challenged is not a method. Stating what the protocol cannot establish is what makes the parts it can establish worth anything.

Where are the results produced by this protocol?

Published with their conditions at norg.ai/case-studies. Pricing is at norg.ai/pricing. Norg Pty Ltd, ABN 44 669 712 494, founded 14 July 2023.