Robin Ooi — AI SEO Malaysia
AI Search & GEO

How Does ChatGPT Decide Which Sources to Cite?

By Robin Ooi6 min read

The short answer

ChatGPT answers most questions by running its own searches, pulling back a ranked list of pages, and writing from the top few. Retrieval position is the single strongest predictor of whether a page gets cited, which makes citation mostly a trust and visibility problem rather than a formatting one.

Diagram of a retrieval graph showing one highly cited source node connected to lower-weighted candidate pages
Retrieval works like a graph. A handful of trusted nodes get selected repeatedly, and everything else sits at the edges.

What actually happens when you ask ChatGPT a question

Most people picture ChatGPT reading the whole internet and picking a favourite. That is not what happens. When a question needs current information, the model does something much closer to what you would do: it turns your question into several searches, collects a ranked set of pages, reads the top ones, and writes an answer from what it found.

That middle step is where everything is decided. The model can only cite what the retrieval layer handed it. If your page never comes back in that ranked set, nothing about your writing matters. You were not rejected. You were never in the room.

This is the part most advice about AI search skips, and it is why so much of that advice does not work.

What the research says about retrieval position

AirOps and Kevin Indig ran the largest public study on this so far, analysing roughly 815,000 query-page pairs across 16,851 ChatGPT queries. The headline finding is blunt: retrieval rank predicts citation better than anything else they measured.

Retrieval rank is the strongest predictor of whether a page gets cited in a ChatGPT answer. A page at rank 0 is 4x more likely to be cited than a page at rank 10.
AirOps, The Fan-Out Effect

In their data, the top retrieved result was cited around 58% of the time. The tenth fell to roughly 14%. That is a steep curve, and it holds regardless of how well written the page at position ten happens to be.

The second finding is the one that upsets people. Pages covering only part of a question sometimes beat pages that tried to cover everything. The instinct to make an article longer and more comprehensive so the AI has more to work with is, on this evidence, backwards. Focus wins.

Retrieval position and citation likelihood
Retrieval positionRoughly how often citedPractical read
1st~58%You are effectively the answer
Mid packFalls steadilyCited when the answer needs breadth
10th~14%Occasional, unreliable, hard to build on
Not retrieved0%Content quality is irrelevant here

So what actually drives retrieval?

Retrieval is a search problem wearing a new coat. The systems that feed AI answers lean on the same signals classic search has always leaned on, filtered through a tighter trust screen. Which means the honest answer to how does ChatGPT choose sources is: mostly the same way search engines decide who is credible, then a preference for whoever sits at the top of that list.

Domain-level trust

Retrieval systems repeatedly favour a small set of domains inside any given topic. Study any set of AI answers in your industry and you will see the same eight or ten sites over and over. That recurring group is your niche trust set. Getting into it is the actual job.

Third-party corroboration

A claim that appears only on your own website is a claim with one witness. When the same expertise shows up on publications, industry roundups, comparison pages and community discussions, the retrieval layer has multiple independent reasons to treat you as a real entity. This is the single most durable lever available, and it happens off your website, which is why it gets skipped.

Query-page fit

Once you are retrievable, precision starts to matter. A page that answers one specific question cleanly outperforms a page that mentions the topic somewhere in the middle of 4,000 words. This is where structure earns its keep, and where schema markup for AI search does real work by labelling what a page is about.

Freshness, where it applies

For anything with a date attached, pricing, tools, regulations, best-of lists, recency is a strong tiebreaker. For evergreen definitions it barely registers. Refresh what genuinely decays and leave the rest alone.

The uncomfortable part of this

There is a question worth asking every time someone pitches you an AI content package. If my content is perfectly optimised but the model still doesn’t trust my domain, what then?

If the answer is vague, or it circles back to more headings and more FAQ blocks, the strategy is incomplete. Formatting is a multiplier on trust you already have. It is not a substitute for trust you do not.

I say this as someone who sells this work. A content-only engagement on an untrusted domain is easy to sell and hard to justify six months later. It is more useful to find out early which of the three states you are in, which is why every AI SEO engagement I run starts with that diagnosis rather than a content calendar.

Three states, three different jobs
StateWhat you seeWhat the work should be
TrustedCited regularly across related questionsExpand coverage, defend the positions, keep pages current
BorderlineShows up sometimes, disappears on rewordingReinforce authority, then sharpen the pages that nearly made it
UntrustedEffectively absent across the whole question setOff-site authority first. Content work here mostly wastes money

What to do about it this quarter

The sequence matters more than any individual tactic. Doing these in the wrong order is how budgets get burned.

  1. Build a question set, not a keyword list. Thirty to fifty questions a real buyer would type, covering problems, comparisons, pricing and objections. This becomes your measurement baseline.
  2. Run them and record who gets cited. Normalise to root domains. You are looking for the recurring names, not the one-offs. This is your niche trust set.
  3. Work out where you sit. Absent, sporadic or present. Be honest about it, because this decides everything downstream.
  4. Close the authority gap before the content gap. Editorial mentions, expert commentary, inclusion in category roundups, original data other people can cite. If branded queries are part of the problem, negative autocomplete suggestions feed the same retrieval layer and need handling alongside.
  5. Then sharpen the pages. Once you are being retrieved, precision and structure convert that retrieval into citations.
  6. Re-run the same question set monthly. Same questions, same day of month. Movement only means something against a fixed baseline.

If you want the mechanics of step six, I wrote a separate piece on how to track AI search visibility without pretending rank tracking still applies. And if you are still working out how this differs from ordinary optimisation, start with GEO vs SEO.

Frequently asked questions

Does ChatGPT read my whole website?

No. When a question needs live information, ChatGPT runs searches and retrieves a ranked set of individual pages, then reads the top few. It sees the pages retrieval hands it, not your site as a whole. This is why one strong page can get cited while the rest of your site stays invisible.

Why does ChatGPT cite my competitor and not me?

Almost always because their pages are retrieved higher than yours for those questions. Retrieval position is the strongest predictor of citation, and it is driven by domain trust, third-party mentions and classic search visibility. It is rarely because their writing is better than yours.

Will adding FAQ sections and schema get me cited?

They help once your pages are already being retrieved, because they make your content easier to parse and attribute. They do very little if your domain sits outside the trusted set for that topic. Structure is a multiplier on visibility you have, not a replacement for visibility you lack.

Do I need to block or allow AI crawlers?

Allow them if you want to be cited. Check your robots.txt for GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot and Google-Extended. Some security products and CDN rules block these by default, which quietly removes you from the running before anything else is even assessed.

How often does ChatGPT update which sources it uses?

Continuously, because the retrieval layer runs live searches at answer time rather than using a fixed index snapshot. That means citation sets can shift week to week, which is why measuring against a fixed question set on a regular schedule is more useful than checking one query occasionally.

Is longer content better for AI citation?

Not automatically. The AirOps and Kevin Indig research found that focused pages covering part of a question sometimes outperform comprehensive pages attempting to cover all of it. Clarity and specificity matter more than word count. A tight 1,200 word answer often beats a padded 3,000 word one.

Robin Ooi, AI SEO and online reputation management specialist based in Penang, Malaysia

Written by Robin Ooi

Robin is a Malaysian AI SEO and reputation specialist with more than fifteen years in search. He is the Amazon bestselling author of Your SEO Sucks! and works with MNC and listed company clients across Kuala Lumpur, Penang, Johor and Singapore. Read more about Robin Ooi.

See exactly where you're losing traffic — in 60 seconds

Run the free AI SEO scan on your website. You'll get your visibility score across Google and AI search, plus the top issues costing you leads right now. No sign-up call, no obligation.

Trusted by MNCs and listed companies across Malaysia & Singapore

WhatsApp Robin