There's no single public algorithm explaining exactly how ChatGPT, Perplexity or Gemini choose which sources to cite — but the observable patterns are consistent across engines.
1. Direct answers, not introductions
Pages that answer the implicit question in the first 100-150 characters get extracted more easily than ones with a long intro before the point. An LLM (like a hurried reader) prefers not to have to "dig".
2. Structured data as a trust shortcut
An Organization or FAQPage JSON-LD block doesn't "convince" an LLM by itself, but it reduces ambiguity: official name, site, entity type become certain facts instead of inferences from text.
3. Consistency across pages (entity consistency)
If the brand name, address or description shift slightly from page to page, an engine struggles to recognize it's the same entity — and when in doubt, it cites the less ambiguous source, even if less complete.
4. Technical accessibility comes first
None of this matters if the engine's bot can't even read your site: a robots.txt that blocks GPTBot or ClaudeBot excludes the site from the competition before it even starts.
Want to know which of these your site meets today? Run a free scan: the report lists them one by one with the fix ready to download.