Most AI visibility tools can produce an attractive percentage, a list of cited sources, and a dashboard full of movement. That does not mean the measurement is reliable enough to...
Most AI visibility tools can produce an attractive percentage, a list of cited sources, and a dashboard full of movement. That does not mean the measurement is reliable enough to guide content, brand, or budget decisions. Before buying, you need to establish whether the tool observes the engines your audience uses, lets you control the prompts being tested, exposes the underlying answers, explains its scores, and can reproduce a test later.
Those are buying criteria, not feature-boxes to tick. A tool that fails any one of them may still look sophisticated while making it impossible to answer a basic question: did visibility change because your brand changed, or because the measurement changed?
Use the following table as a screening document during demos and trials. Treat “disqualifying answer” literally. If a supplier gives that answer, do not assume the limitation will be fixed after purchase.
| Criterion | Why it matters | Disqualifying answer |
|---|---|---|
| Engine coverage | Different AI answer engines retrieve, cite, summarize, and personalize information differently. A result in one engine is not a proxy for performance everywhere. | “We cover the major engines,” without naming them, showing the exact access method, or confirming which models and interfaces are tested. |
| Prompt-set control | You need to test real commercial questions, branded questions, category questions, and competitor comparisons. A fixed vendor-generated set can hide the queries that matter to your business. | You cannot upload, edit, tag, pause, or delete prompts, or the supplier will not show the full prompt list before purchase. |
| Raw answer export | Scores alone cannot reveal whether a citation is prominent, whether the answer is factually wrong, or whether your brand is mentioned in a damaging context. | There is no export of the complete answer text, citations, timestamp, and prompt for every observation. |
| Scoring transparency | A visibility score is only useful when you know how mentions, citations, rank, sentiment, and answer completeness are weighted. | The score is described as proprietary and the supplier cannot provide the formula, component metrics, or a worked example. |
| Re-run comparability | Trend data must distinguish genuine change from changes in prompts, model versions, sampling, location, or settings. | The tool silently replaces prompts, changes test conditions, or cannot show the methodology and timestamp for each run. |
| Operational usability | Insights need to reach content, PR, product, and leadership teams in a format they can act on within a normal reporting cycle. | There is no scheduled export, API, role-based access, or way to assign an observation to an owner. |
“AI visibility” is not one channel. A tool may test a public chatbot, a search-generated answer, an API model, or a simulated response. Those are materially different environments. They can use different retrieval systems, geographic assumptions, knowledge cutoffs, citation behavior, and safety filters.
Ask the vendor to name every engine and interface included in your plan. Then ask five follow-up questions:
For an initial programme, set a minimum coverage requirement of three materially different environments: one conversational assistant, one search-integrated answer experience, and one additional engine that matters to your customers. This is a purchasing threshold, not a claim that three engines represent the whole market. If your audience is concentrated in one country or language, location and language support may matter more than adding a fourth engine.
Do not pay extra for a long list of model names if the tool cannot tell you whether it is testing the public interface your customers use. “We monitor 20 models” can be less useful than transparent coverage of four relevant experiences.
The prompt set is the measurement instrument. If you cannot inspect and control it, you do not control the research.
A credible setup should let you import prompts from a spreadsheet, edit them in bulk, tag them by intent, and preserve their history. Useful tags include brand, category, problem, comparison, location, audience, product, and funnel stage. It should also support negative or exclusion terms where necessary, such as prompts that apply only to a specific market.
Start with at least 50 prompts for a small test and 100 or more for a recurring programme. Fewer than 30 prompts can be useful for a smoke test, but it is too narrow for a headline visibility score. A practical first set might contain:
These numbers are a planning framework, not a universal sample-size law. The important requirement is that every prompt has a business reason and an owner. Remove prompts that never influence a decision, but do not let the vendor quietly substitute easier queries because your original set produces weak results.
Popular advice is often to “track a handful of high-value prompts.” That advice is wrong when it turns five hand-picked questions into a company-wide visibility KPI. High-value prompts belong in a focused segment, not as the entire measurement universe.
A chart showing that your brand was mentioned in 42 percent of tests is not evidence you can audit. You need the observation behind the chart.
At minimum, each exported record should include the exact prompt, full answer, cited sources, engine, model or interface, location, language, timestamp, run identifier, and any relevant system settings. Ideally, it should also include whether your brand was mentioned, where it appeared, which competitors appeared, and the tool’s classification of sentiment or recommendation.
Test the export before signing. Ask for a sample containing at least 100 observations and check whether:
A raw export is also a governance safeguard. AI answers can contain outdated prices, incorrect product claims, invented associations, or competitor comparisons that require a response. A score cannot tell a product or communications team what needs correcting. The answer itself can.
Every AI visibility tool needs to reduce observations into summaries. The danger begins when the summary is presented as an objective fact rather than a model with choices.
Ask the supplier to define each component. For example, a score might combine mention rate, citation rate, answer prominence, sentiment, and share against named competitors. Those components are not interchangeable. A brand mentioned once at the end of an answer is not equivalent to a brand recommended in the first sentence. A citation to an outdated page is not necessarily a valuable citation.
Request a worked calculation using five real or sample answers. You should be able to see:
Set a governance rule that no executive KPI uses a composite score until its components have been documented. A safer reporting pattern is to show three to five separate measures, such as unprompted mention rate, citation rate, recommendation rate, and factual accuracy. Add a composite index only if its weights reflect a stated business objective.
Do not choose the tool with the highest-looking score. Scores from two platforms are not comparable unless they use the same prompts, engines, sampling, definitions, and weighting. A 60 on one platform may represent a different phenomenon from a 60 on another.
AI outputs are variable. The same prompt may produce a different answer on another run because the model sampled differently, retrieved different pages, or changed its underlying system. Variation is not automatically a measurement failure, but a vendor must make it visible.
During a trial, run the same 20 prompts three times within one day and again after seven days. Compare the raw answers and the summary metrics. Ask whether the tool uses one run, multiple samples, or a stored answer. If it samples multiple times, ask how many: a single response may be cheap but noisy, while 5 or 10 responses provide a better view of variation at a higher cost.
Require a fixed test protocol for recurring measurement. It should specify prompt wording, engine, location, language, account state, browsing setting, run frequency, sample count, and model version where available. A useful monthly programme might run 100 to 300 prompts with the same settings, then flag any methodology change alongside the results.
Look for change logs. If the engine changes model, the tool changes its classifier, or a prompt is edited, the historical chart should mark a break in the series. A smooth line that conceals these events is more dangerous than a volatile line that explains them.
Pricing varies by prompt volume, engine access, seats, API use, historical retention, and whether human analysis is included. As an internal buying framework, reserve a trial budget of approximately $500 to $2,000 for a short evaluation if the supplier charges for one. For a recurring programme, compare the total annual cost against the number of usable observations, not the number of dashboard seats.
Ask for the price at three volumes: 100 prompts, 500 prompts, and 2,000 prompts per month. Confirm whether every engine consumes a separate credit and whether re-runs consume additional credits. Also check whether raw exports, historical data beyond 90 days, API access, and additional users are charged separately.
A low monthly price is not a saving if the data cannot be exported or repeated. Conversely, a high price is not justified by a large prompt allowance if the prompts are fixed and the score is opaque. The economic unit that matters is a comparable, auditable observation your team can act on.
Use a two-week evaluation rather than a guided demo alone. Give each shortlisted provider the same 100 prompts, the same competitor list, and the same target markets. Score the trial against the following checks:
Reject a trial if the vendor supplies only screenshots, a hand-picked success report, or a single polished score. Also reject any contract that allows the provider to alter your prompt set or methodology without notification. Data that cannot be inspected, repeated, and challenged is not a dependable basis for an SEO or brand decision.
The best AI visibility tools are not necessarily the ones with the most alerts, the largest model count, or the most impressive dashboard. They are the ones that make measurement boringly consistent: the same prompts, known engines, complete answers, explicit scoring, and visible changes in method.
Before purchase, write the acceptance criteria into the contract. Set minimum engine coverage, a required prompt-management workflow, raw-export fields, score documentation, re-run rules, retention period, and change-notification terms. If a supplier cannot meet those conditions, buy less—or do not buy—rather than treating an attractive but unauditable number as evidence of visibility.
Want the measurement, not the pitch?
Send us your domain. We run the baseline on your category prompts and send back the raw answers alongside the score — you can check our working.
Đội ngũ chuyên gia Vidco Group sẵn sàng đồng hành cùng bạn
Bước 1 / 4
Chúng tôi sẽ liên hệ trong vòng 2 giờ làm việc.