SEO visibility was a settled question. You had Search Console for impressions, Ahrefs or Semrush for positions, GA4 for what happened after the click, and three tools that disagreed by a few percentage points about numbers that were broadly the same.
AI visibility has none of that. There is no console that tells you how often ChatGPT named you. There is no position to report. Ask the same question twice and you can get two different answers, with a different set of brands in each. Every vendor selling an AI visibility score is using their own prompt set, their own sampling frequency, and their own definition of what counts as an appearance, which is why two tools can look at the same brand and give you scores thirty points apart.
Nobody has standardised this yet. That's inconvenient, and it's also the reason getting it right is worth doing now: you're building an instrument while your competitors are still waiting for one to be handed to them.
Start With the Traffic You Can Actually See
There is a real, measurable slice of this, and it's the sensible place to begin.
AI assistants send referral traffic. When someone clicks a link inside a ChatGPT answer, that visit lands in GA4 like any other referral. It just doesn't get labelled usefully, because GA4's default channel groupings don't know what an AI assistant is and dump most of it into Referral or Unassigned.
Fix that first. In GA4, go to Admin, Channel Groups, copy the default group to create a new one, and build a channel that catches the AI sources by hostname: chatgpt.com, perplexity.ai, claude.ai, gemini.google.com, copilot.microsoft.com, and openai.com. Ahrefs has a step-by-step walkthrough if you want the screenshots.
Once that's live you can segment AI sessions from everything else and look at them properly: which pages they land on, how long they stay, what they convert at.
You'll also notice the numbers don't add up. Referral data is systematically undercounted here. Some assistants strip the referrer, some open links in embedded browsers, and some users read the answer and then type your domain in directly a day later, which lands in Direct with no trace of where the idea came from.
Ahrefs' own research puts LLM traffic at roughly 0.1% of visits while acknowledging the figure is almost certainly far too low because platforms withhold source data. Whatever GA4 tells you, assume the real influence is larger and unmeasured.
Citations and Mentions Are Not the Same Metric
This distinction is the one most teams get wrong, and it changes what you build.
A citation is when an AI answer links to your site as a source. It's attributable, it can generate a click, and it shows up in your referral data.
A mention is when the model names your brand in the text of its answer without linking anywhere. No click, no referral, no trace in any analytics tool you own.
Both influence buying decisions, and in B2B the mention is often the more valuable one. When a buyer asks for the best platforms in your category and the model returns four names in a paragraph, being one of those four names is the entire outcome. Nobody clicks. The shortlist just forms.
This is why a traffic dashboard alone will mislead you. It measures the citation side and is structurally blind to the mention side, which is where most of the value sits. You need a second data source, and no one will give it to you. You have to generate it.
Building the Dashboard
The core mechanic is simple: ask the models the questions your buyers ask, on a schedule, and record what comes back. Everything else is structure around that.
1. Build a prompt set. Start with 30 to 50 questions and grow toward 100 or more. Cover the full spread: category questions ("best X for Y"), comparison questions ("X vs Y"), alternative questions ("alternatives to [competitor]"), problem questions phrased the way a buyer would phrase them, and branded questions about your own company. Tag each one by buyer stage and product line so you can slice results later.
2. Run them on a fixed cadence, with repeats. Weekly is the right starting rhythm. Run each prompt three to five times per cycle, because a single response is noise. Model output varies run to run, and if you only sample once you'll spend your Monday explaining a drop that isn't real.
3. Record these fields for every response. Whether your brand appeared. Where in the answer it appeared, since first mention carries more weight than a passing reference at the bottom. Which competitors appeared alongside you. Whether the mention was a recommendation or just a name in a list. The sentiment.
4. Compute a small number of metrics. Presence rate, or the percentage of responses that mention you. Share of voice against your named competitors. Citation source mix. Sentiment and accuracy rate. Resist the urge to build twenty metrics. Four that move decisions beat twenty that decorate a slide.
5. Layer in the click-side data. Pull GA4 AI sessions and conversions, Search Console impressions, and, if you use Bing at all, the AI Performance Report in Bing Webmaster Tools, which is currently the only first-party citation reporting any platform offers. Google Search Console still merges AI Overviews into total search performance with no way to separate them.
6. Make one view per audience. An executive wants presence rate and share of voice over time. A content lead needs the prompt-level table showing which questions you lose and who wins them. Build both from the same data or nobody will look at either.
Turning the Data Into a Loop
A dashboard that only reports is a waste of everyone's Tuesday. The point is the feedback cycle, and it runs like this.

Read the losses first. Sort your prompt table by the questions where you're absent and a competitor is present. That list is your content roadmap, already prioritised by buyer intent. For each one, look at which sources the model cited instead of you and go earn a presence in those specific places.
Mine the answers for new prompts. Every response you collect contains vocabulary you didn't write. The way models phrase your category, the adjacent concerns they raise unprompted, the comparison sets they assume. Feed that language back into your prompt set each cycle. Your query bank should grow every month, and it should grow from what the models actually say rather than from a keyword tool that was never built for conversational queries.
Fix misrepresentation at the source. When a model puts you in the wrong category or quotes stale pricing, publishing a correction on your own blog rarely moves it. Your citation source mix will tell you which third-party page the wrong information is coming from. Update that, or get a more credible source to state the right version, then re-measure in 30 days.
Re-run the same prompts after every change. This is the part teams skip. Ship the fix, note the date, and check the same prompts four weeks later. AI visibility work has a long and uneven lag, and without a before-and-after on a fixed prompt set you have no idea whether anything you did worked.
Tracking, content, and correction feed each other here. Run it for a quarter and you'll have something no tool can sell you: a record of what actually moves visibility for your brand specifically, rather than a score somebody else computed with a methodology they won't show you.




