AI visibility is not one universal rank. The same brand can appear in one answer, disappear in another model, and show up only when a question is phrased more specifically. That makes the test design more important than a headline score.
Build a question set from buyer decisions
Start with 10–20 questions a real buyer could ask before choosing a product. Use a mix of jobs:
- Category discovery: “What tools help a small SaaS track AI recommendations?”
- Problem diagnosis: “Why do competitors appear in AI answers but my brand does not?”
- Comparison: “Which AI visibility tool fits a small SaaS team?”
- Use case: “How can I track prompts that mention my brand?”
- Trust: “How can I verify where an AI answer got its product information?”
Do not fill the list with small wording variations. Each question should represent a different decision or objection.
Keep the first test controlled
| Field | What to save |
|---|---|
| Date and time | When the check happened |
| Surface | ChatGPT, Gemini, or Perplexity |
| Exact question | Copy it; do not summarize |
| Brand mentioned | Yes / No |
| Brand recommended | Yes / No / Ambiguous |
| Context | Why the brand was included or excluded |
| Competitors | Other products named |
| Sources | URLs shown in or attached to the answer |
| Evidence | Screenshot or stored response, subject to platform rules |
Use a clean, documented test environment where possible. The goal is not to pretend variables disappear. It is to record enough context that the next check is comparable.
Separate mentions, recommendations, and citations
Mention
The brand name appeared.
Recommendation
The product was presented as a viable choice for the job.
Citation
A page was used or shown as a source.
A brand can be cited without being recommended, or recommended without a visible citation. Track the signals separately. Save negative or uncertain wording too; a high mention count can hide weak positioning.
Connect the baseline to one product action
Visibility data becomes useful only when it changes a decision. For each weak question, inspect the cited pages and ask:
- Do we have a page that answers this exact job?
- Does that page state who the product is for and what it produces?
- Does it include a verifiable sample, method, price, or limitation?
- Is the page discoverable and internally linked?
- Do third-party sources describe us consistently?
Choose one change for the next cycle. Record the page edited and the date, then rerun the same question set. If visibility changes, treat it as a signal for investigation—not automatic proof of causation.
Manual checks or continuous monitoring?
Manual checks are enough for a first baseline or a very small question set. They become difficult when you add products, competitors, languages, models, or weekly comparisons. A monitoring product should save the exact prompt set, responses, competitor pressure, sources, and trend so the team can spend its time on the next action.
Geo Agent turns a one-time audit into a product workspace: save buyer questions, monitor selected models, review competitors and sources, choose actions, and recheck. Start with the free audit if you do not yet have a baseline.
What this baseline can and cannot prove
A controlled prompt set creates a repeatable proxy signal. It does not create a universal market-share score, remove model randomness, or prove that one page edit caused a later mention. Preserve the underlying question-by-surface evidence even when you summarize results.
Frequently asked questions
How many questions should the first baseline use?
Ten to twenty distinct buyer questions are enough to expose patterns without creating an unmanageable project. Start with questions closest to product selection.
Should brand-name prompts be included?
Yes, but keep branded and non-branded questions separate. Branded questions test whether public facts are understood; non-branded questions test discovery before the buyer chooses you.
Can results from different models be averaged into one score?
They can be summarized, but preserve the underlying results. One average can hide that one surface recommends you while another never mentions you.
What is the first useful outcome?
A saved baseline and one specific action. The commercial outcome comes later when a visitor starts the product experience, reaches checkout, or pays. Visibility alone is not revenue.