There are established ways to measure your website's position on Google: rankings, clicks, impressions. AI assistants are different. There is no results page, every user sees a different answer, and answers change over time. In this article we explain what to pay attention to if you want to measure AI visibility in a meaningful way.
What does AI visibility mean?
AI visibility is how often your brand is mentioned in the questions your buyers ask an AI assistant. Think of questions like "What are the best accounting software tools?" or "Recommend a company in Ankara that builds corporate websites." If your name appears in the answer, you are visible. If it does not, the buyer may never hear of you at that moment.
In Turkey, Google's AI Mode and AI Overviews have been available since February 18, 2026. So in Turkey, these answers can now reach not only ChatGPT users but anyone searching on Google.
Why is asking once not enough?
AI answers are unstable. In SparkToro's research, 600 volunteers ran the same 12 prompts through ChatGPT, Claude and Google's AI features a total of 2,961 times. The chance of two answers containing the same list of brands was under 1%.
The same research reaches two important conclusions:
- Looking at individual answers, or asking "what position am I in on AI," is misleading.
- According to the researchers, a brand's mention rate across many prompts, each asked more than once, is a reasonable indicator.
This gives the basic rule of measurement: instead of looking at a few questions once, ask a fixed list of questions at regular intervals.
Which questions should you ask?
The value of a measurement depends on the quality of the question list. A good question list follows these points:
- Your brand name does not appear in the questions. If you ask "Is company X any good?", the engine will mention you anyway; that measures awareness, not visibility.
- The questions are sentences a buyer would actually type. Not a pile of keywords, but a natural request: "Which ERP would you recommend for a small manufacturing company?"
- Different intents are covered: category questions ("best X"), comparison questions, local questions ("X in Izmir"), and problem-based questions ("What should I use for problem X?").
- The list is kept fixed. If you change the questions every week, you cannot compare the results.
In AIShortlist, the question list is suggested from your site, and the suggestions do not include your brand name. No measurement starts until you edit and approve the list. The reason for this step is simple: the person who knows the business best should choose the questions.
What does a sample question list look like?
Imagine a hypothetical agency in Izmir, Turkey, that builds corporate websites. Its buyers might type the following into an assistant:
- Which companies in Izmir build corporate websites?
- What should I look for when getting a website built for a small manufacturing company?
- For a corporate website, is a ready-made template or a custom design better?
- Can you recommend a reliable agency in Izmir that builds e-commerce sites?
- I want to redesign my website. Which agency should I work with?
The agency's name does not appear anywhere in this list. The questions cover different intents: local search, information search, comparison and a direct request for a recommendation. For information questions, the engine usually does not mention brands; it still helps to have them on the list, because they show which sources are being consulted.
When you prepare your own list, ask your sales team and customer service: what do buyers ask most often in the first conversation? Those sentences are usually very close to the sentences people type into AI.
Which engines should you measure?
Each engine should be measured separately. ChatGPT, Gemini and Perplexity look at different sources and produce different answers. You can be strong in one and invisible in another. Merging the engines into a single average hides this difference. Today AIShortlist measures ChatGPT, Gemini and Perplexity; Google AI Mode and AI Overviews are not included in the measurement.
When an engine's answer cannot be retrieved, that result should be recorded as missing. Filling the gap with another engine's answer corrupts the measurement.
Which metrics should you look at?
Four metrics are enough to read a measurement result:
- Visibility score: the share of answers in which your brand is mentioned. This is the main metric.
- Share of voice: your share of all brand mentions in the answers. It shows where you stand next to competitors.
- Rank among brands: where you rank among all the brands the engines mention.
- Average position: what position you appear in on the list when you are mentioned.
A note of caution on position: the research above shows that position within a single answer is very unstable. So it is better to read position as a trend over weeks, not on its own. The real decision metrics are the visibility score and share of voice.
Why should you track competitors too?
Your visibility score alone does not say whether things are good or bad. If a competitor is mentioned many times more often than you on the same questions, the gap is large. If no brand in the category is mentioned often, a score that looks low may actually be average.
That is why the measurement should also record the other brands that appear in the answers. Over time, a "you vs. competitors" chart shows whether your work is closing the gap.
How often should you measure?
Two different things should not be mixed up here. Individual answers are volatile: the same question can give two different lists on the same day. A brand's overall standing changes slowly: your mention rate does not change on the day you fix a page or get onto a list page. That is why weekly measurement is a reasonable balance for most brands. More frequent measurement helps in fast moving sectors, or when you want to watch a few critical questions closely; remember that every measurement is a paid call.
When you read the rate, look at the denominator too. 10 questions and 3 engines make 30 answers; if you are named in 15, the rate is 50%. If an engine did not answer that day, the denominator shrinks and the rate rests on fewer answers. How AIShortlist does this calculation is on the methodology page.
Keep the day and time of measurement as fixed as you can. When you change a page or get onto a new list page, note the date. When you look at the chart a few weeks later, these notes let you see what happened after which change.
What do you do after measuring?
Measurement is not a goal in itself. The result should answer two questions:
- Which pages do the engines read when they answer? If "best X" list pages are among the cited pages and you are not on those lists, that is where the first job is.
- Is your own site easy for AI to read? Are bots blocked, do your pages clearly say what you do, are there sections that answer questions directly?
The answers to these two questions also define what you expect to change in the next measurement.
Should you measure by hand or with a tool?
Measuring by hand is a good exercise for getting started with a few questions. Open a spreadsheet, put the questions in rows and the engines in columns, and note in each cell whether your brand was mentioned. Asking more than 20 questions on three engines every week turns that spreadsheet into a job of several hours each week.
When choosing a tool, check whether the engines are measured separately, who decides the question list, and whether the results come with the source pages. You can try AIShortlist for free: 1 brand, 10 questions, one measurement across three engines and a technical site check, with no credit card.
One more note: no measurement tool guarantees that AI will recommend you. Measurement lets you see honestly what is working and what is not.