To track AI search visibility, run a fixed set of prompts that mirror real buyer questions across ChatGPT, Perplexity, Gemini and Google’s AI features on a schedule, and log whether you’re mentioned, cited, how prominently and how accurately. Pair that with a GA4 channel for AI referrals and a monthly share-of-voice report against named competitors.
Why AI visibility needs its own measurement
Rank tracking assumes one query, one results page and ten positions. AI answers break each of those assumptions:
- No fixed positions. An assistant may name you first, fifth or only as a linked source. Position means prominence, not a rank number.
- Answers vary. Ask the same question twice and you can get different brands, so you need repeated runs and rates.
- Mentions and clicks are decoupled. A buyer can read your name in ChatGPT, then search your brand on Google. GA4 records that as branded organic or direct, not AI.
- Some AI surfaces are invisible in analytics. Clicks from Google AI Overviews and AI Mode arrive as ordinary Google organic traffic, and Search Console folds them into regular web search totals.
So you need two instruments: an answer-side panel (what the engines say) and a traffic-side channel (what visits arrive). Neither is complete; together they show direction. Still deciding whether AI search deserves budget at all? Start with GEO vs SEO.
Building a prompt set that mirrors buyer questions
The prompt set is your panel. Build it from what buyers actually ask, not what you wish they asked, or every metric downstream will flatter you.
Where to source prompts
- Sales call recordings, for the exact words prospects use to describe the problem
- Win/loss notes that name the alternatives buyers considered
- Long, question-shaped queries in Search Console
- Support tickets and onboarding questions
Rewrite each one the way a buyer types into a chat box, with context. “We’re a 30-person UK agency, what’s the best project management tool for client approvals?” is a prompt. “Project management software” is a keyword.
Cover five prompt types
| Type | Example | What it tells you |
|---|---|---|
| Category | “Best inventory software for small wholesalers” | Whether you’re in the consideration set at all |
| Problem | “How do I stop stockouts across two warehouses?” | Whether engines connect your product to the pain |
| Comparison | “[Competitor] vs [you]”, “Alternatives to [competitor]” | How you’re framed next to rivals |
| Use case | “Inventory tool for a Shopify brand using a 3PL” | Fit for your best-fit segments |
| Brand | “Is [you] good for wholesalers?”, “[You] pricing” | Accuracy and sentiment about you |
Size, engines and runs
For most B2B and ecommerce brands, 30 to 80 prompts is a workable range. Weight it toward category and use-case prompts, where buying decisions start, and keep brand prompts to about a fifth of the set.
A sensible default engine list is ChatGPT, Google’s AI Overviews and AI Mode, and Perplexity. Add Gemini, Copilot or Claude if they come up in sales calls. AI Overviews trigger on searches, so check the search-query version of each prompt and note whether an overview appears at all.
Because answers vary, run each prompt at least three times per engine per cycle and record the share of runs that mention you. Use clean sessions: logged out where possible, memory off, consistent location. Log the date and model, because model updates can shift results overnight.
Freeze a core set for trend reporting and keep new prompts in a separate exploratory tab.
Tracking mentions, citations and sentiment
Log one row per prompt run. These are the fields I use:
| Field | What to record |
|---|---|
| Mentioned | Yes or no: is your brand named in the answer? |
| Position | First recommendation, one of a list, or passing mention |
| Cited | Is your domain linked as a source, and which URL? |
| Competitors named | Every tracked competitor that appears |
| Sources cited | Third-party domains the engine linked |
| Accuracy | Wrong facts about pricing, features, markets or company details |
| Sentiment | Positive, neutral or negative, plus the phrase that sets it |
Keep mentions and citations separate, because they fail for different reasons. Mentioned but never cited often means the model knows you but finds no page worth quoting. Cited but not recommended often means useful content and a weak third-party reputation.
The “sources cited” column is the most underrated. Across a few dozen prompts, the same review sites, community threads and listicles keep appearing. That list is your target map for reviews, PR and partnerships.
Give accuracy its own flag. When I review AI answers for a brand, outdated pricing, retired plans and the wrong target market are common, and usually fixable at the source.
Measuring AI referral traffic in GA4
Answer-side tracking shows visibility. GA4 shows whether it sends visits that convert.
In the default channel group, AI assistant visits end up split across Referral, Unassigned and Direct. A custom channel group gives them a channel of their own:
- In GA4, go to Admin > Data display > Channel groups and copy the default group to create a new one.
- Add a channel called “AI Assistants” with the condition Source matches regex and a pattern such as the one below.
- Move the new channel above Referral. GA4 assigns each session to the first channel whose rules match, so if Referral sits higher it absorbs the AI traffic.
- Save, then switch your Traffic acquisition reports and explorations to the new group.
.*(chatgpt\.com|chat\.openai\.com|perplexity\.ai|gemini\.google\.com|copilot\.microsoft\.com|claude\.ai|chat\.deepseek\.com|meta\.ai).*
Then report sessions, key events and landing pages for the channel. Landing pages show which content engines send people to.
Three caveats:
- It’s a floor. Many AI app visits arrive without a referrer and count as Direct, and some buyers never click, searching your brand later instead. Watch branded impressions in Search Console and add “AI assistant” to your “How did you hear about us?” field.
- Google’s AI surfaces aren’t separable. AI Overview and AI Mode clicks show as google / organic.
- The regex goes stale. Review the domain list every quarter.
Manual tracking vs monitoring tools
Manual tracking means a spreadsheet and discipline. A worked example with made-up round numbers: 40 prompts across 3 engines with 3 runs each is 360 answers a month. At a minute or two per answer, that’s 6 to 12 hours of work.
Monitoring tools, both standalone AI visibility platforms and AI modules inside SEO suites, run prompts on a schedule and chart share of voice for you.
| Manual spreadsheet | Monitoring tool | |
|---|---|---|
| Best for | Up to about 50 prompts, the first few cycles | Larger sets, several markets, weekly cadence |
| Cost | Team time | Subscription, usually scaling with prompts and engines |
| Strength | You read every answer, so you catch nuance and wrong facts | Volume, repeat runs, trend charts |
| Weakness | Doesn’t scale; inconsistent if several people run it | May query an API instead of the consumer app, so results can differ from what buyers see |
Before you buy, ask the vendor:
- Do you query the consumer interfaces with web search on, or the API?
- How many runs per prompt, and do you report a rate or a single result?
- Can I export the raw answers and cited sources, not just scores?
My usual approach: track manually for two cycles to learn what good data looks like, then move the frozen core set into a tool and keep reading a sample of raw answers monthly.
Reporting AI share of voice
AI share of voice = your brand mentions ÷ total mentions of all tracked brands across the prompt set.
Continuing the example with made-up numbers: across 360 answers, you’re mentioned in 72, Competitor A in 150, Competitor B in 110 and Competitor C in 38. That’s 370 brand mentions in total, so your share of voice is 72 ÷ 370, about 19.5%. Your mention rate is 72 ÷ 360, or 20%.
Report both. Share of voice is relative and can rise when a competitor drops out even if you did nothing.
The monthly one-page report:
- Share of voice and mention rate, overall and by engine, with a three-month trend
- Mention rate by prompt type: category, problem, comparison, use case, brand
- Citation rate and your most-cited URLs
- Top third-party domains cited, and whether you’re present on them
- Accuracy issues found and fixed
- The AI Assistants channel in GA4: sessions, key events, top landing pages
- Actions taken and actions planned
Read trends over a quarter, not a month. A model update can move numbers without anything changing on your side, so annotate known model releases.
Turning findings into actions
Each pattern points to a different fix:
| Finding | Likely cause | Action |
|---|---|---|
| Absent from category prompts while competitors appear | Thin presence in the sources engines cite | Target the review sites, lists and communities in your “sources cited” column |
| Mentioned but rarely cited | No page answers the question directly | Publish pages with clear, quotable answers, comparisons and specs |
| Wrong pricing or features | Stale pages, listings or articles | Fix your own site first, then the profiles engines cite |
| Negative framing | A specific review, thread or article | Trace the source; respond, fix the issue or earn newer coverage |
| Visible in Perplexity, missing in ChatGPT | Crawler access or index coverage | Confirm OAI-SearchBot isn’t blocked in robots.txt or at your CDN, and key pages are indexed in Bing |
| AI referrals that don’t convert | Landing pages without a next step | Add a relevant call to action toward demo, trial or product pages |
Measurement tells you where to aim; the work is content, reputation and technical access. In my GEO and AI search visibility work, I set up the tracking, run the monthly cycle and ship the fixes rather than hand over a report.
Get it built
If you want to know whether AI assistants recommend you, and what would change that, I can build the prompt set, set up the GA4 channel and run your baseline. Start with the fixed-price Growth Audit, credited if we continue. See pricing or get in touch.