“Are we showing up in AI answers?” is a reasonable question with an unreasonable number of bad answers attached to it. Most of what gets sold as AI visibility measurement...
“Are we showing up in AI answers?” is a reasonable question with an unreasonable number of bad answers attached to it. Most of what gets sold as AI visibility measurement fails one of three tests: the prompt set is not fixed, the engine coverage is not stated, or the scoring collapses several different things into one number. Here is a method that passes all three, written so you can run it yourself.
The method below is engine-agnostic, but every figure we quote from our own runs comes from a generative engine API, measured on a stated date. It is not a measurement of Google’s AI Overviews, which requires a different collection method entirely. Any provider who shows you one number covering “AI search” without naming the engines behind it is showing you an average of things that do not average.
The single most common error is putting the brand name in the prompt. Asking a model “what do you think of Acme Corp?” measures whether the model has heard of Acme. It tells you nothing about whether a buyer who has never heard of Acme would be shown it.
A usable prompt set has these properties:
Run every prompt, store the full text of every answer, and record the date and engine version. Storing raw text is not bureaucratic overhead — it is the only thing that lets you or a third party re-score the run later under a different definition, or audit a number you find surprising.
Two collection details that change results more than people expect: temperature, and whether the engine’s own web-browsing or grounding is enabled. A grounded run and an ungrounded run answer meaningfully different questions. Fix both settings and record them.
| Metric | Definition | What a low score means |
|---|---|---|
| Mention Rate | Share of answers in which your brand name appears. | The model does not associate your brand with the category at all. This is an entity and third-party-coverage problem. |
| Citation Rate | Share of answers in which your domain appears as a source. | Your content is not being treated as a citable source, even where your brand is known. This is a content-format and authority problem. |
| Share of Model | Your mentions divided by all vendor mentions across the run. | You are present but losing the comparison. This is a positioning problem, not a volume problem. |
Keep them apart. A brand can sit at a healthy Mention Rate and a Citation Rate of zero — that combination is common, and it prescribes completely different work from the reverse.
This is where automated scoring quietly breaks. Naive substring matching produces false positives at an embarrassing rate: a two- or three-letter brand token will match inside longer unrelated words, and in Vietnamese it will also match across diacritics. Match on word boundaries, normalise accents on both sides of the comparison, and hand-check a sample of at least twenty hits before you trust any aggregate. We have seen a careless matcher inflate a mention count by a factor of three.
The same run that scores you also tells you who the model currently treats as the category. Count every vendor named, rank them, and keep the ranking. This is usually the most commercially useful artefact of the whole exercise, and it costs nothing extra to produce.
Re-run at a fixed interval — 90 days is a sensible default — with the identical prompt set, identical engine settings, and identical scoring code. Movement is then attributable to your work rather than to a changed test. If you must extend the prompt set, keep the original as a frozen core and report it separately.
It will not tell you traffic, and it should not pretend to. Mention Rate is a visibility measure, not a demand measure. It also cannot separate a genuine ranking improvement from a model update that happened to reshuffle its associations — which is exactly why the competitor board matters: if every vendor moved, the model changed, not your site.
Fifty is the practical floor. Below that, one unusual answer moves the headline percentage by two points and you will over-read noise. One hundred is comfortable.
Whichever your buyers type. If they type both, run both and score them as separate sets — results routinely differ, sometimes substantially.
It is normal for a brand with little third-party coverage, and it is useful, because it is unambiguous. Our own first baseline was 0% on both mention and citation.
Yes, and that is the point of publishing it. Ask any provider for their prompt set, their engine list, their dates and their raw answers. All four, or the number means nothing.
Want your own numbers instead of ours?
Send us your domain. We run the baseline on your category prompts and send back the raw answers, not just a score.
Đội ngũ chuyên gia Vidco Group sẵn sàng đồng hành cùng bạn
Bước 1 / 4
Chúng tôi sẽ liên hệ trong vòng 2 giờ làm việc.