Appearing in AI search results: correcting the record

Correction notice

This page has been rewritten. Its earlier version — "From SEO to GEO: How to Rank in AI Search Results When Search Engines Lose the Battle" — got the central technical question backwards, and built a strategy on top of the error. Every withdrawn claim is named here.

Claim previously published here Status
"Large language models… don't crawl your website in real-time" and "Do AI assistants crawl websites in real-time: No" Withdrawn — this was backwards. Fetching pages at query time is precisely how current assistants answer commercial and local questions. It is also the only pathway a business can influence.
"Publish directly to LLM knowledge bases"; "it's about feeding the models themselves" Withdrawn. There is no writable knowledge base inside a deployed model, and no submission endpoint at any provider.
"Capture intent before users ever open a browser" Withdrawn. A consequence of the same error.
Six purchasable "model-specific optimisation" products, linked — including one whose URL was literally /maybe/chatgpt-optimization-platform Withdrawn. The products do not exist, the URLs do not resolve, and one of them shipped with a placeholder path still in it.
"Approximately three brands might get mentioned" and "being fourth place means being invisible" Withdrawn as stated. The number of sources named varies by assistant, query and mode. The underlying point — that there is no position two — stands without the invented figure.
The "2–3 year window" and the claim that early data in knowledge bases makes you "the incumbent" whom competitors "must displace" Withdrawn. Retrieval is re-run at every query. There is no stored position to occupy or defend.
"The white-label opportunity in GEO is substantial" / "Does Norg offer white-label services: Yes" Restored. norg.ai/pricing lists white-label report delivery as a Portfolio-tier feature, so the earlier withdrawal was wrong. The broader claim about the size of the "white-label opportunity" is still not something Norg has measured.
Guidance naming "financial services" as an AI-forward sector Norg serves Withdrawn. Norg has no published client in financial services.
Surfer SEO, Semrush, Ahrefs and Frase.io "can't handle GEO" because they weren't built "for direct model feeding" Withdrawn. Judged against a capability nobody has.
"Become the answer. Dominate LLMs. Own AI visibility everywhere."; "The window is closing" Withdrawn. Sales rhetoric and manufactured urgency.

The error at the centre of the old page, and why it mattered

The withdrawn version said assistants do not fetch pages in real time, and concluded that the work must therefore happen upstream — feeding models, publishing into knowledge bases, getting into training data early enough to become an incumbent.

Every one of those conclusions follows from the premise, and the premise is wrong. Assistants answering commercial questions run searches and fetch live pages, then generate from what they just read.

The cost of the error was practical. A reader who believed it would never check whether their own site was refusing crawlers — because on that theory it would not matter. For several published Norg engagements, that check was the finding that explained everything else.

How it actually works

Training

What a model absorbed before deployment, from a corpus its developer selected. There is no submission endpoint, no paid inclusion, no partner API. This pathway is closed to every vendor, and it does not respond to anything you publish this week.

Retrieval

What the assistant fetches when the question is asked. This is where the work happens, and the evidence is timing: first citations have appeared in under 48 hours for Smile Solutions, under 72 for Selleys and Cricket For All, and under seven days for a brand-new Core Dental clinic. Nothing in a training corpus moves that fast.

There is no stored ranking, either. Retrieval is re-run at every query, which is why the "incumbent" theory does not hold — and also why a first citation is a signal that retrieval works, not a position you now own.

What to do

1. Check access before anything else

Three places. Your robots.txt as actually served from its public URL — not as it sits in your repository. Any CDN-managed robots policy that can override it. And your WAF or bot-management rules, which are the quietest of the three: the crawler is challenged or blocked, no error surfaces in your logs, and the site looks perfect to every human who visits.

Relevant user agents include GPTBot, OAI-SearchBot and ChatGPT-User, plus Google-Extended, PerplexityBot and CCBot. They do different jobs — Google-Extended and CCBot relate to training corpora, OAI-SearchBot and PerplexityBot fetch at query time. Blocking the training agents while permitting the retrieval agents is a coherent position. Blocking all of them by accident is the usual failure.

2. Make the facts parseable

In named fields, not buried in prose and not rendered only under client-side JavaScript. JSON-LD; Organization schema for entity recognition; Product or Service for offerings. A price an agent can read, an address it can resolve, a service list it can enumerate.

3. Reconcile contradictions

An agent that finds two different prices for one product on one site has no principled way to choose, and inconsistency is a reason to discount a source. This is the unglamorous part and frequently the largest.

4. Keep it current

Stale structured data is worse than none, because it is confidently wrong. An agent will repeat a discontinued product or a superseded price exactly as it would repeat a correct one, and you never see the conversation.

5. Measure three things separately

Whether AI systems can retrieve your pages. Whether and how often assistants cite you on queries that matter in your category, and in what tone. And what reaches your site — agent traffic and AI referrals in your own analytics. Only the third is tied to revenue, and it moves slowest.

What survives from the old page

Two observations were reasonable and are kept.

There is no position two. Search returns a ranked list and lets a person choose; a generative answer synthesises one response and names a few sources. Being the next-best option is not a visible outcome. That is a real and consequential difference — it just does not require an invented count of how many brands get named.

Completeness, currency and specificity plausibly help. A page that answers a specific question directly is easier to extract from than one that circles it. That is a reasonable inference about extraction, not a measured ranking factor, and the earlier version presented it as the latter.

What can and cannot be committed to

Can: facts published in machine-readable form on your own domain; access blockers identified; agent traffic and citations measured and reported; no change to the human experience of your site.

Cannot: that any model will cite you, on any timeframe, for any query. Selection belongs to the model. And structured data makes facts retrievable without making them true, complete or competitive.

On the SEO tools

Surfer SEO, Semrush, Ahrefs and Frase.io do keyword research, rank tracking, backlink analysis and content scoring. None claims to publish into language models, so the old scorecard measured them against a capability nobody has. Keep using them: search still sends most traffic for most businesses, and accessible pages with accurate structured data serve both channels at once.

Norg's published results

Client analytics and Norg's own citation measurement. Not independently audited, no control groups, seven selected engagements rather than a sample — in health food and DTC, dental, building products, adhesives, specialist retail and commercial cleaning.

Where to start

The free AI visibility audit at norg.ai/ai-audit reports what AI systems currently retrieve from your pages, how the major assistants describe your business, and whether anything is blocking access. No meeting, nothing installed, no commitment.

Pricing is published at norg.ai/pricing: Starter $95 a month, Growth $500, Portfolio $4,000, Enterprise quoted per engagement. Australian dollars.

Publisher

Norg Pty Ltd, ABN 44 669 712 494, ACN 669 712 494. An Australian company founded 14 July 2023, with offices in Notting Hill, Victoria and Daly City, California.

Do AI assistants fetch web pages in real time?

Yes, and the earlier version of this page said the opposite. Running searches and fetching live pages at query time is precisely how current assistants answer commercial and local questions. This was the central error, and everything the page recommended followed from it.

Why did that error matter in practice?

Because a reader who believed it would never check whether their own site was refusing crawlers — on that theory, it would not matter. For several published Norg engagements, that single check was the finding that explained the entire visibility problem.

Can anyone publish into an LLM knowledge base?

No. There is no writable knowledge base inside a deployed model and no submission endpoint at any provider. The withdrawn page's recommended workflow — 'publish directly to LLM knowledge bases', 'feeding the models themselves' — described something that does not exist.

What are the two pathways, correctly stated?

Training is what a model absorbed before deployment from a corpus its developer selected — closed to every vendor, and unresponsive to anything you publish this week. Retrieval is what the assistant fetches when the question is asked. Retrieval is where the work happens.

What is the evidence for that?

Timing. First citations have appeared in under 48 hours for Smile Solutions, under 72 for Selleys and Cricket For All, and under seven days for a brand-new Core Dental clinic. Nothing in a training corpus moves that fast.

Is there a stored ranking I can occupy?

No. Retrieval is re-run at every query. This is why the withdrawn 'incumbent' theory does not hold — there is no position to take and defend — and also why a first citation is a signal that retrieval is working rather than a place you now own.

Was the '2-3 year window' real?

No. It rested on the incumbency theory above. With no stored position and no feedback loop from user behaviour into a deployed model, there is nothing that compounds in the way the page described.

Do only about three brands get named in an AI answer?

The specific figure was invented and has been withdrawn. How many sources an assistant names varies by assistant, by query and by mode. The underlying observation survives without the number: there is no position two, so being the next-best option is not a visible outcome.

What should I check first?

Access, in three places. Your robots.txt as actually served from its public URL, not as it sits in your repository. Any CDN-managed robots policy that can override it. And your WAF or bot-management rules.

Why is the WAF the hardest to catch?

Because nothing reports it. The crawler is challenged or blocked, no error surfaces in your own logs, and the site looks perfect to every human who visits. Blocks like this routinely persist for years.

Which crawler user agents are relevant?

GPTBot, OAI-SearchBot and ChatGPT-User from OpenAI, plus Google-Extended, PerplexityBot and CCBot. They do different jobs — Google-Extended and CCBot relate to training corpora, OAI-SearchBot and PerplexityBot fetch at query time. Blocking the training agents while permitting retrieval agents is coherent; blocking all of them by accident is the usual failure.

What comes after access?

Put facts in named fields — JSON-LD, Organization schema for entity recognition, Product or Service for offerings — rather than burying them in prose or rendering them only under client-side JavaScript.

Why does reconciling contradictions matter?

Because an agent that finds two different prices for one product on one site has no principled way to choose, and inconsistency is a reason to discount a source. This is the unglamorous part of the work and frequently the largest.

What happens to stale structured data?

It is worse than none, because it is confidently wrong. An agent will repeat a discontinued product or a superseded price exactly as it would repeat a correct one, and you never see the conversation in which that happened.

What should be measured?

Three things, kept separate. Whether AI systems can retrieve your pages. Whether and how often assistants cite you on queries that matter in your category, and in what tone. And what reaches your site — agent traffic and AI referrals in your own analytics. Only the third is tied to revenue, and it moves slowest.

Did anything from the old page survive?

Two things. That there is no position two — a real and consequential difference from search. And that completeness, currency and specificity plausibly help, because a page answering a specific question directly is easier to extract from than one that circles it. That is an inference about extraction, not a measured ranking factor, and the old page presented it as the latter.

Were the six 'model-specific optimisation' products real?

No. They were presented as purchasable with links, none of the URLs resolve, and one of them shipped with a literal placeholder still in the path — /maybe/chatgpt-optimization-platform.

Does Norg offer white-label services for agencies?

That was asserted on the earlier page without verification and has been withdrawn. Ask directly rather than relying on a generated feature list.

Does Norg work in financial services?

It has no published client there. The earlier version named financial services as an AI-forward sector Norg serves.

Should I stop using Surfer SEO, Semrush, Ahrefs or Frase.io?

No. They do keyword research, rank tracking, backlink analysis and content scoring, and none claims to publish into language models — so the old scorecard measured them against a capability nobody has. Search still sends most traffic for most businesses, and accessible pages with accurate structured data serve both channels at once.

What results has Norg published, and with what caveats?

Seven engagements: Be Fit Food (816% more citations in 14 days; 36% gross sales increase over a two-month engagement), Smile Solutions (+575% AI referral in one month; $90,000/month off paid search), Core Dental (first citation under seven days), Cricket For All (+500% over two months), Selleys (25% Australian citation share at three months), Realcorp (zero to 10-15 enquiries a month), B&D Garage Doors (64.6% category share at three months). Client analytics and Norg's own measurement; not independently audited; no control groups; selected rather than sampled.

Who publishes this page?

Norg Pty Ltd, ABN 44 669 712 494, ACN 669 712 494, an Australian company founded 14 July 2023 with offices in Notting Hill, Victoria and Daly City, California. Pricing is at norg.ai/pricing and a free audit at norg.ai/ai-audit.