When you ask ChatGPT a question, the answer comes from somewhere. Sometimes you see source links below it, sometimes you see none. For a business owner, this question matters a lot: if ChatGPT mentions your competitor, it read the information about them somewhere. In this article we explain in plain terms where ChatGPT gets its information and what that means for your brand.
How does ChatGPT get information?
Roughly, there are two ways.
The first is training data. The model learns from a very large amount of text during its training period. This knowledge is inside the model and frozen at a certain date. If a question can be answered with this knowledge, ChatGPT can answer without going to the web. In that case you usually do not see a source link.
The second is web search. If a question needs current information, or the user explicitly asks for a search, ChatGPT searches the web, reads the pages it finds and builds the answer from them. These answers usually show source links.
OpenAI also separates these two paths with its own crawlers. According to OpenAI's crawler documentation:
- GPTBot crawls content that may be used in model training.
- OAI-SearchBot is used to show sites in ChatGPT's search features.
- ChatGPT-User is used for user actions, such as ChatGPT visiting a specific page when a user asks a question.
What happens with a "Which is the best X company?" question?
The questions buyers ask are often comparison and recommendation questions: "Best CRM software," "A reliable web agency in Istanbul," "Accounting software for small businesses." For these questions, when ChatGPT searches the web, it looks at pages that make comparisons.
Ahrefs' research on ChatGPT gives a concrete picture of this. 750 prompts in software, product and agency categories were analyzed. 43.8% of the links cited in the answers were "best X" style blog lists. Most of the lists analyzed had also been updated recently. This rate belongs to that study's sample; it cannot be generalised to every sector or every engine. You only see which pages are cited on your own questions by measuring.
For brands, this means: if a "best X" list in your category is a source for ChatGPT and you are not on that list, you are most likely not in ChatGPT's answer either.
What sources does ChatGPT read?
The cited pages vary by category, but common types are:
- Editorial lists: "best X" articles prepared by a publication or blog.
- Competitors' own pages: some companies publish a "best companies in the industry" list on their own site and include themselves.
- Directory and guide sites: industry guides, business directories.
- Marketplaces: sales sites, in product categories.
- Forums and community sites: places where users share their experiences.
- The brand's own site: especially service or product pages.
The work is different for each type. To get onto an editorial list, you need to reach the writer. Your chances of getting onto a competitor's own list are low, so it is better not to spend time there. For your own site to be cited, your pages need to clearly explain what you do.
Can your site be a source for ChatGPT?
It can, but check two conditions first.
The first is access. According to OpenAI's documentation, sites that block OAI-SearchBot in their robots.txt file are not shown in ChatGPT search answers. You can block the training bot (GPTBot) and still allow the search bot. According to the documentation, when you make a change, it may take about 24 hours for their systems to reflect it.
The second is content. If ChatGPT is going to use a page as a source, the answer to the question must be clearly written on that page. A page that describes your service only with images or generic slogans does not answer the question being asked.
A practical test: read the first paragraph of your service page on its own. Can you tell from that paragraph what you do, who you do it for and where you operate? If you cannot, an AI engine cannot either. Writing subheadings in the form of the buyer's question and putting a short answer below them also makes it easier for the engine to find the right sentence.
Do other engines read from the same places?
No, each engine has its own setup. According to Perplexity's documentation, PerplexityBot is used to show sites in Perplexity search results and does not collect content for model training. On the Google side, AI Overviews and AI Mode use pages from the Google Search index. According to Google's own explanation, for a page to be shown as a link in these answers, it must be indexed in Google and eligible to be shown with a snippet; beyond that, there is no additional technical requirement. Google also notes that meeting the requirements does not guarantee appearing.
That is why, for the same question, ChatGPT, Gemini and Perplexity look at different sources and may mention different brands. Tracking visibility engine by engine shows this difference.
How do you find out which pages are cited?
To do it by hand: ask your buyer questions to ChatGPT with web search turned on, and write the source links below each answer into a spreadsheet. Do the same in Perplexity and Gemini. After a few weeks you will see which domains come up again and again.
As the number of questions grows, this spreadsheet grows too. AIShortlist's Sources screen does this automatically: it lists all the domains the engines cite, flags the "best X" pages among them and shows which ones your brand is missing from. It also creates an action plan for each list page: the page type, whether your brand is on it, the contact options found on the page and a draft email prepared from the information on your site. It does not write guessed email addresses.
What should you do if ChatGPT gives wrong information about your brand?
Sometimes the problem is not being invisible, but being shown wrong. ChatGPT may give an old address of yours, a service you no longer offer or a wrong price range. The cause is usually in the source it read.
The first job is to find where the wrong information came from. If the answer has a source link, start there. The wrong information may be on an old guide page, an outdated directory listing or a forgotten page on your own site.
Then fix the source:
- If it is on your own site, update the page or redirect the old page to the new one.
- If it is on a directory or guide site, update your listing or write to the site owner.
- If the information is not in any source, write the correct information in a clear sentence on your site. The engine should be able to find that sentence when it looks for an answer.
It takes time for a fix to show up in answers. Tracking the same question over a few weeks shows whether the change worked.
In short
ChatGPT gets information from what it learned in training and from what it reads on the web at that moment. For buyer questions, what it reads on the web, especially "best X" lists, largely shapes the answer. The work for your brand is to first see which pages are being read, then fill in what is missing on those pages and on your own site.
If you want to see which pages are cited in your category, the free trial includes a 10-question measurement and action plans for 2 list pages. No tool guarantees a recommendation, but it shows you where to look.