How AI chooses sources: what we know, what we measure, and what we're just guessing
A generative system typically first searches for or retrieves relevant material, then composes an answer from it. The exact weighting of sources isn't public and varies by service. That's why it isn't honest to present a single universal list of "GEO ranking factors" as a verified algorithm.
A simplified process
- The system interprets the query and its context.
- It retrieves candidate documents or data.
- It evaluates their relevance and usefulness.
- It generates an answer from the selected material.
- Depending on the product, it displays citations, links, or related sources.
A document used during generation isn't always explicitly cited, and a cited link doesn't necessarily prove it was the only source used.
What Google confirms
Google states that AI Overviews and AI Mode are built on its Search index and core ranking and quality systems. They may use query fan-out — that is, multiple related searches.
From this, we can safely conclude:
- a page must be indexable and eligible to appear;
- technical SEO remains relevant;
- useful original content has a lasting advantage;
- neither llms.txt nor special AI schema is required.
What we cannot conclude from this is the precise weight of an H2 heading, paragraph length, or the number of citations for an AI Overview.
What OpenAI confirms
OpenAI states:
- a public website can be included in ChatGPT Search;
- not blocking OAI-SearchBot is important for summaries and citations;
- ranking is based on multiple factors focused on relevant and trustworthy information;
- top placement cannot be guaranteed;
- GPTBot and OAI-SearchBot serve different purposes.
OpenAI does not publicly disclose a complete list of ranking factors or their weights.
What research shows
An academic GEO paper tested nine content modifications on a specific benchmark and measured the resulting change in visibility within generative answers. Adding citations and statistics produced strong results; keyword stuffing didn't help.
The study is valuable evidence, but:
- it measures a specific experimental setup;
- it doesn't describe today's internal algorithm for every product;
- the percentages can't automatically be promised to every website;
- the effect varied by topic and starting position.
What's reasonable practice, though not guaranteed
Accessibility
A system can't reliably use content it can't reach.
Direct relevance
A page should answer a specific need and include the necessary detail.
Verifiability
Primary sources, methodology, an author, and current data reduce risk.
Original value
First-hand experience and data set a site apart from many identical summaries.
Unambiguous entities
A consistent name, author profile, organization, and product data reduce ambiguity.
Reputation
Independent, high-quality mentions can confirm that a brand genuinely belongs in a topic.
These points make sense for people and for search alike, but their exact weight in any specific answer isn't known.
How to run your own test
- Pick a single page and a set of prompts.
- Record baseline answers across multiple repetitions.
- Make one significant change.
- Allow enough time for crawling and indexing.
- Repeat the same set under the same conditions.
- Track mentions, citations, sources, and variability.
- Treat the result as an observation, not a universal law.
Myths
- AI always cites the top three organic results. Query fan-out and product differences can bring in other sources.
- Schema guarantees a citation. No.
- llms.txt is a ranking factor. Not for Google; for other systems, no universal effect has been documented.
- The longer the article, the better. Length without value doesn't help.
- More mentions always mean higher authority. Quality, context, and authenticity are what matter.