Skip to content

How does AIShortlist measure AI visibility?

What the numbers represent, how each engine is asked, the score formulas and the limits. This page describes how the product works today.

Last updated: 6 October 2026. If the way the product measures changes, this page is updated too.

What do the numbers represent?

The results in the dashboard represent the answers each engine's official API gives to your approved question list, with web search on. Each question is asked to each engine once per measurement.

This is not an exact copy of the answer a buyer sees in the app. Account, chat history, location and the model in use at that moment can make the app's answer different. Read the numbers as an indicator you can compare week to week under the same conditions, not as an absolute truth.

How is each engine asked?

EngineWhere it is askedWeb searchWhere sources come from
ChatGPTOpenAI's official API, a model that searches the webOnLink citations in the API response
GeminiGoogle's official Gemini APIOn (grounding with Google Search)Grounding sources in the API response
PerplexityPerplexity's official APIOnSearch results and citations in the API response

We do not change sampling settings such as temperature; the provider's default applies. As providers update their models, the model used for measurement can change too, which can affect comparisons with earlier weeks.

Each engine is measured separately. If one fails, another is never used in its place.

How is the question sent?

The question sent is the sentence you approved. A short instruction, the same for every engine, is added to the end: rely on current web sources, name specific brands, answer as a numbered list and gather the companies you named on one line at the end.

The instruction does two things: answers become comparable across engines, and the list of named brands can be extracted reliably. The instruction is in Turkish for Turkish questions and in English for all others.

There is no system prompt and no chat history. Each question is asked on its own, in a clean session.

Limit: a buyer types the same question into the app without this instruction. The shape of the answer and the brands it names can differ for that reason.

How are language and location handled?

No separate country, city or language setting is sent to the engines. The language is the language of the question itself. Location matters when it is written in the question, as in "recommend an agency in Istanbul for AI visibility".

The language and city you set on a question set decide the language of the questions we suggest and how many of them are local. If you want to measure for a specific country or city, write the place into the question.

How many times is a question asked?

In one measurement each question is asked to each engine once. AI answers can change from one attempt to the next even for the same question; a single measurement is one sample.

So look at the trend over weeks, not at a single week. If a question already has an answer from the last 24 hours it is not asked again; that answer is used.

How are mentions, position and competitors extracted?

No second AI interprets the answer; extraction is done by text matching.

  • Mention: you count as named if your brand name, or one of the other spellings you entered on the question set, appears in the answer.
  • Position: your place in the list of companies the engine named. Your brand can appear in the answer without being in the list; then it counts as named but has no position.
  • Three tiers: "Recommended" when your brand is in the engine's list of suggested companies, "Known" when your name is in the answer but not in that list, "Absent" when it is not there at all. Recommended and Known both count as mentioned in the score.
  • Competitors: the other names in the same list. Company suffixes (Ltd, Inc, GmbH and similar) are dropped and different spellings of the same brand are merged.
  • Quote: the sentence around the place your brand appears, shown in the dashboard under "Where you were mentioned".

Limit: matching depends on spelling. If your brand name is very short or a common word, false matches can happen; if the answer spells it differently, it can be missed. Adding the other spellings in the question set settings reduces this.

How are the scores calculated?

MetricHow it is calculated
AI visibility scoreAnswers that name your brand, divided by answers measured. Each question and engine pair is one answer: 10 questions and 3 engines make 30 answers.
Share of voiceYour mentions, divided by your mentions plus every competitor mention in the answers.
Rank among brandsYou and the 10 most mentioned competitors are ordered by mention count; this is your place in that order.
Average positionThe average of your positions in the answers where you were named and had a position.
Visibility by enginePer engine: answers from that engine that name you, divided by answers received from that engine.

Scores are calculated from the latest full measurement. Questions tracked daily are kept separate and do not mix into the weekly score.

What happens if an engine does not answer?

If a call fails it is tried twice more, 20 and 40 seconds apart. If there is still no answer, that question counts as not measured on that engine; the cell shows "Not measured" in the question table.

No zero is written for a missing answer and no other engine is used. A missing answer is left out of the denominator: the score is calculated only from answers that were really received.

If an engine gives no answers or only some, the overview, the weekly email and the report say so separately: you see which engine answered how many questions.

Limit: a week with missing answers rests on fewer answers. When you compare two weeks, check whether either one carries a missing answer notice.

What does the Sources screen show, and what does it not?

Sources are the links the engine returns with its API response; they are not guessed from the answer text. Tracking parameters are removed, the same address is counted once and one answer counts a domain at most once.

These links show which pages the engine looked at when it built that answer. On their own they do not prove why a brand was recommended.

What do we store, and what do you see?

The full text of every answer, its sources and the extracted list of companies are stored. In the dashboard you see, for every question and engine, whether you were named, your position and the sentence where your brand appears.

The full answer text is shown in the dashboard on the Answers screen: what each engine wrote for every question in the last measurement, with its cited sources.

How often is it measured?

On paid plans every question set is measured automatically every 7 days. The free trial runs one measurement. The dashboard and the report show the date of the latest measurement.

On Growth, Pro and Agency you mark the questions that matter most (5, 15 and 40 questions); these are tracked every day and shown in a separate table.

What happens to history when the question list changes?

A question is never deleted, it is closed; its past results stay. If you edit a question, the old one is closed and the new one starts as a separate question. After every change you approve the list again; nothing is measured without approval.

The trend chart adds up every answer measured on a given day. If you changed the question list substantially, before and after do not measure the same thing; note the date of the change.

When you mark a job as done on the to-do list, the state of the related questions on that day is saved and the next measurement is shown beside it. This is an observation; because AI answers also change on their own, it is not proof that the job caused the change.

How are rival labels and notifications produced?

The rival radar and the notifications are not a new measurement: the stored results of the last 12 measurements are compared with each other. No AI call is made.

A company's share is its mentions in that measurement divided by your mentions plus the mentions of every rival.

  • Passed you: its share rose above yours and was not above it in the measurement before.
  • New: named for the first time in the last 12 measurements. Back: missing in the measurement before, present in earlier ones.
  • Rising and Falling: its share moved by at least 1 percentage point, or its best position in the answers changed.
  • Gone: named in the measurement before, not named in this one.
  • Engine notification: your mention rate on one engine moved by at least 10 percentage points, or that engine started or stopped naming you.
  • Question notification: lost when at least one engine named you and now none does; won when the opposite happened.
  • New source: a site never shown in the earlier measurements was shown as a source in this one.

Limit: a label tells what differs between two measurements, not why. AI answers also change on their own, so one label is not a trend. The timeline covers the last 12 measurements.

How is question intent decided?

The intent label in the question portfolio comes from the wording of the question, by fixed word rules; no AI is used. Checked in this order: Brand when your brand name is in the question; Alternative for words like "alternative", "instead of", "other than"; Comparison for "which one", "difference", "compare", "vs"; Price when price or cost is asked; Local when a city or "near me" appears; Discovery otherwise.

The label only groups the list. It does not change the measurement, the score or the order of the questions.

How is the strategy summary written?

A text model writes the strategy summary, but it does not search the web and gets no outside information. It is given only the results of that measurement the panel already shows, as plain sentences: recommended, known and absent counts per engine, questions won and lost, questions where you are absent, the pages shown as sources, the priority findings of the site audit.

Every step it writes is checked by code. A step containing a number or a site name that is not in those sentences is dropped. A step containing a guarantee or an expected score gain is dropped. The summary is written once per measurement and stored.

Limit: the summary orders and words what was measured; it is not a new measurement. Doing a step does not guarantee that you will be recommended; the next measurement shows what follows.

How are fact accuracy and site consistency checked?

In the fact accuracy check each engine is asked about your business with one question: your brand name, domain and city if any, and the phone, email, address, opening hours, founding year and services. Web search is on; the company list instruction used for buyer questions is not added to this question.

A text model extracts the value stated for each field and quotes the sentence that contains it word for word. A value whose quote is not found in the answer is not counted. The verdict is not left to the model; it is made by code.

  • Phone: the last 10 digits are compared; country code and leading zero do not count as a difference.
  • Email and founding year: correct when the value you entered appears in what the engine said, wrong when it said another value.
  • Address: correct when at least 60 percent of the words of your address appear in what the engine said.
  • Opening hours and services: free text, so no correct or wrong verdict; "Stated" is shown when the engine gave a value.
  • Does not know: the engine gave no value for that field. It is not counted as wrong and is left out of the accuracy rate. The rate is correct divided by correct plus wrong.
  • Site consistency: your home page and up to two contact or about pages are read; the name, phone, email, address and year you entered are looked for in the plain text of the pages and in phone links. JavaScript is not run.

Limit: each engine is asked one question once; the same engine can answer differently on another day. In the consistency check "not found" does not mean the information is missing: it may be inside an image or on a page we did not read.

How are the site crawl and bot access tested?

  • Crawl: under the name AIShortlistBot, following robots.txt rules and the links on your site, up to your plan's page limit.
  • Speed: measured with Google PageSpeed Insights, mobile and desktop.
  • Bot access: three checks. The robots.txt rules are read; a trial request carrying the bot's name is sent to your home page (OAI-SearchBot, ChatGPT-User, GPTBot, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Googlebot, Bingbot); the raw HTML is checked for readable text.

Limit: the trial request comes from our server. Real bots are also recognised by IP address, so this result is shown as a "possible block" and is not definitive. The definitive evidence is in your own server logs.

What do we not measure?

  • Google AI Mode and AI Overviews. The Gemini measurement does not stand in for them.
  • Claude, Copilot and other assistants. The engines measured today are ChatGPT, Gemini and Perplexity.
  • The personalised answer a logged-in user sees in the app.
  • Your site traffic and sales. The dashboard does not connect to Google Analytics or Search Console.
  • A guarantee of being recommended. We measure and show the steps that make it more likely; the result depends on your content and on third parties.

Related pages

See where you stand today

Create a question set, approve your questions and get your first result. Free, no card needed.

Chat on WhatsApp