In June 2026 Google added a new set of reports to Search Console. They show how often pages from your site appeared in AI Overviews, AI Mode and Discover's AI features. By 31 August they were live for every site.
It's useful data. But look at what it counts: your URLs, shown as sources. It doesn't tell you whether the answer named your company, what it said about you, or which competitors it recommended instead.
That gap is what AI share of voice fills.
I wrote this for marketing leads and founders at established B2B companies who've been asked "are we showing up in ChatGPT?" and want a number they can defend. It covers the five metrics worth tracking, what the new Google and Bing reports can and can't do, how many times you need to ask each question, and a scorecard you can copy into a spreadsheet this afternoon.
What AI Share of Voice Means, and How It Differs From Search
AI share of voice counts who gets named in AI answers. In search, share of voice usually means your slice of visibility for a set of keywords, worked out from rankings and estimated clicks. An AI answer has no list of ten links to rank in. The assistant writes a paragraph and names a handful of companies, or none at all.
Google's AI Overviews are the AI-written summaries that appear above some Google results. Since January 2026 Gemini 3 has been the default model behind them, and people can ask a follow-up question and carry on into AI Mode. ChatGPT, Gemini, Perplexity and Claude write answers in a similar way, so the same measurement works across all of them.
Why bother? Because AI summaries change what people do next. Pew Research Center tracked the real browsing of 900 US adults in March 2025. About 18% of their Google searches produced an AI summary. When one appeared, people were more likely to stop browsing altogether.
| Search page with an AI summary | 26% |
|---|---|
| Search page with only traditional results | 16% |
Source: Pew Research Center, browsing data of 900 US adults, March 2025, published July 2025. US data. UK behaviour may differ. Chart by ThriveFinity.
If a buyer reads the answer and stops, being named in that answer is the visibility that counts. A page-one ranking they never scroll to doesn't help much.
One more distinction, because it trips people up. AI share of voice is a comparative number. It only means something next to the competitors you chose to track. If you pick weak competitors, you'll look great. Pick the companies you actually lose deals to. If you're not sure who they are, my guide to competitive intelligence for B2B companies shows how to find out.
The Five Numbers to Track
Five measures cover what most B2B companies need. Define them once, write the definitions down, and don't change them halfway through a year.
| Metric | Formula | What it tells you |
|---|---|---|
| Mention rate | Answers that name you ÷ all answers | How often you appear at all |
| AI share of voice | Your mentions ÷ mentions of all tracked brands | Your slice of the conversation against competitors |
| Recommendation share | Answers that recommend you ÷ answers that recommend any brand | Whether you're suggested or just listed |
| Citation share | Citations to your pages ÷ all citations | Whether AI tools use your site as a source |
| Accurate-mention rate | Accurate mentions ÷ all your mentions | Whether what's said about you is right |
Mention rate and share of voice can move in opposite directions, and that's worth understanding. Say a new competitor starts appearing in lots of answers. Your mention rate might hold steady while your share of voice drops, because the pie got bigger. Neither number is wrong. They answer different questions.
Recommendation share is the one I'd watch most closely for B2B. Being listed fifth in "companies that offer X" is nice. Being the one the answer suggests for "a 100-person firm like mine" is what puts you on a shortlist.
Citation share is the odd one out, because it measures your pages, not your name. An answer can cite your blog and recommend someone else. It can also recommend you while citing a directory. Track it, but don't confuse it with being chosen.
Semrush's explainer is a decent visual walk-through of the share of voice idea, and it ends with three ways to improve it, ordered from least to most effort. It's a vendor video built around their own toolkit, so treat the tool part as one option among several.
What Google and Bing Now Report, and What They Don't
Both Google and Microsoft now give site owners some first-party AI data, and you should switch both on. Neither gives you share of voice.
Google's Search Generative AI performance reports launched on 3 June 2026. They show "how often URLs from your site appeared in generative AI features in Search and Discover", broken down by pages, countries, devices and dates. The announcement post was updated to say the reports reached all websites worldwide on 31 August 2026.
Microsoft got there first. In February 2026 it launched AI Performance in Bing Webmaster Tools, covering Copilot, AI summaries in Bing and "select partner integrations". It reports total citations, average cited pages, page-level citation activity and grounding queries, which are the phrases the AI used when it went looking for content to cite.
Aleyda Solis, the SEO consultant, summed up the Search Console version when she finally got access in July:
I finally got access to the new Search Console Generative AI performance report… and: it's easy to feel underwhelmed, especially when compared to the one of Bing or Clarity AI Search dashboards
In the same post she notes it has no prompts, topics or clicks. Just impressions and URLs. Bing's version goes further, because grounding queries show you the sub-searches an AI ran before it cited you. This episode walks through what that data looks like and how to read it:
Here's how I'd use the platform data. It tells you which of your pages AI features pick up, and roughly how often. That's handy for spotting which pages to protect and which to improve. What it can't tell you is whether your company was named when a buyer asked "who should we hire for X?". A page can be cited in an answer that recommends a competitor. Your brand can be recommended in an answer that cites nobody's site at all.
So keep both. Platform reports for your pages. Your own prompt set for your brand.
Building a Prompt Set Buyers Actually Use
The prompt set decides everything that follows. A set built around your own brand name will flatter you. A set of vague industry questions will tell you nothing. Aim for 20 to 50 questions in four groups:
- Category: "Best [category] providers for a [company type] in the UK".
- Problem: the problem in the buyer's own words, before they know a category exists.
- Comparison: "[You] vs [competitor]" and "alternatives to [competitor]".
- Brand: "What does [you] do?", "Is [you] any good?", "How much does [you] cost?".
Take the wording from places buyers really speak: sales call notes, your Search Console queries, People Also Ask boxes and autocomplete. Guesswork from the marketing team tends to produce questions nobody types. If you want a starting list, the free AI Visibility Self-Audit worksheet has twelve buyer prompts you can adapt.
Keep brand questions to a minority. They inflate your score, because an assistant asked about you by name will nearly always mention you.
It also helps to know that a single question isn't really a single search. Google's documentation says AI Overviews and AI Mode may use a "query fan-out" technique, "issuing multiple related searches across subtopics and data sources" to build one response. Bing's grounding queries show the same thing from the other side. So when you write your set, include the sub-questions too: pricing, reviews, "for a company our size", "in the UK".
Once the set is agreed, freeze it. If you want to add questions later, add them as a new group and report it separately. Otherwise this quarter's number can't be compared with last quarter's, and the whole exercise loses its point.
Fix the settings as well as the questions. Write down, for each tool, whether you use the app or the API, whether you're logged in, which country and language you're in, and which model or mode is selected. Answers can differ by country, so test from where your buyers are. Change any of those settings and you've started a new series.
How Many Runs Is Enough?
More than one, and probably more than you'd like. AI answers change from run to run, so a single answer per question is an anecdote.
The clearest public evidence comes from SparkToro's January 2026 study. Six hundred volunteers ran 12 prompts through ChatGPT, Claude and Google's AI a combined 2,961 times. The same list of brands came back less than once in 100 runs. The same list in the same order was closer to once in 1,000.
The study also found something more useful. Some brands showed up again and again even while the order shuffled. One hospital appeared in 69 of 71 answers to its prompt. So the stable thing to measure is how often you appear across many runs, a rate, and the unstable thing is where you sat in any one answer. For accurate tracking, Rand Fishkin's advice in the write-up is blunt. You need to ask "over and over again", usually at least 60 to 100 times, and then average the results.
That's a consumer-scale recommendation, and most B2B companies won't run every prompt 60 times a month by hand. My advice is to pick a number you can sustain, run every prompt that many times in every tool, and never change it. Then treat small moves as noise until they hold for two or three months. A fixed method with modest runs beats a heroic one-off.
It helps to know how big a real change has to be before you believe it. Take the scorecard setup further down: 30 prompts, 4 tools, 3 runs, 360 answers. One extra answer naming you moves your mention rate by less than a third of a percentage point. A swing of five answers, which can easily happen by chance, moves it by about 1.4 points. So if your mention rate goes from 20% to 21.5% in a month, I wouldn't tell the board. If it goes from 20% to 30% and stays there, that's worth a conversation.
Rand explains the thinking behind this in a May 2026 interview, including why "AI rankings" are built on probability and what to track instead:
Measuring Mentions, Position and Citations
For every answer, you're recording three separate things: who was named, in what order, and which pages were cited. Here's the full method:
- Build a fixed prompt set. Write 20 to 50 questions real buyers ask, across category, problem, comparison and brand questions. Keep the list fixed so results can be compared month to month.
- Choose engines and settings. Pick the AI tools your buyers use, such as Google AI Overviews, ChatGPT, Gemini, Perplexity and Claude. Record country, logged-in state and app or API, and keep them the same each run.
- Run each prompt several times. AI answers vary between runs, so run each prompt more than once per engine and record every answer verbatim with its date.
- Extract mentions and citations. For each answer, list every brand named, its position, whether it was recommended, and which web pages were cited as sources.
- Score sentiment and accuracy. Score each mention of your brand positive, neutral or negative, and accurate, outdated or wrong, using a written rubric.
- Calculate and compare. Calculate mention rate, share of voice, recommendation share, citation share and accurate-mention rate for you and each named competitor, and compare against last month.
Position still matters within an answer, even if it's unstable across runs. Being named first in a list of five isn't the same as being fifth. For comparison prompts, also note whether the answer ends with a recommendation and for whom.
Record citations separately from mentions. Where an assistant searches the web, it usually shows its sources. OpenAI's help page says ChatGPT search responses "may include citations", and that ChatGPT "may search the web automatically when your question would benefit from current information." Those cited pages tell you which third-party sites you should be present on.
A caution, though. Rand Fishkin has argued that citation data can mislead you about why a brand gets recommended:
AI tools often already know which brands they're going to search for and recommend, so looking at the sources or URLs ain't gonna help you much.
That's a view, not settled fact, and plenty of people find citation patterns useful. But it's a good reason to keep citation share and share of voice as separate columns. If the model already "knows" the brands it'll name, the work that moves share of voice is mostly about being widely and consistently described across the web, long before anyone types a prompt.
Brand Sentiment and Accuracy in AI Answers: How to Score Them
Score sentiment and accuracy as two separate columns, with a written rule for each label. Sentiment in AI answers is usually mild, so a five-point scale mostly adds noise. Three labels are enough:
- Positive: the answer recommends you, or names a specific strength.
- Neutral: you're listed or described without judgement.
- Negative: the answer names a weakness or a risk, or steers the buyer elsewhere.
Then score accuracy on its own: accurate, outdated (an old price, product or location) or wrong. A positive but wrong mention is still a problem. The buyer will find out, usually on the first call.
Don't assume the answers are right just because they sound confident. Google's own help page for AI Overviews says plainly that they "can and will make mistakes", and labels AI responses as possibly including them. For your brand, the only fair test is to check each statement against your own records.
An AI model can do the first pass of labelling across hundreds of answers, and it'll save you hours. Check a sample by hand every month, keep the rubric fixed, and don't let the same tool both write the answers and mark them without a person looking.
Accuracy is also the metric you can influence most directly. Wrong answers usually trace back to unclear or inconsistent sources: an old pricing page that's still indexed, a stale directory profile, a description that differs from one site to the next. I go through how to fix those at the source in AI visibility for B2B brands.
A Monthly AI Visibility Scorecard (Template)
Use one row per brand and one tab per month. The figures below are illustrative, there to show the arithmetic. They aren't client data. The setup is 30 prompts, 4 AI tools and 3 runs each, which makes 360 answers.
| Brand | Answers naming brand | Mention rate | Share of voice | Recommended | Accurate mentions |
|---|---|---|---|---|---|
| Your company | 72 | 20% | 15% | 18 | 58 of 72 (81%) |
| Competitor A | 198 | 55% | 41% | 95 | n/a |
| Competitor B | 126 | 35% | 26% | 40 | n/a |
| Others | 84 | n/a | 18% | n/a | n/a |
Here's how the numbers fit together. In total, 480 brand mentions were counted across the 360 answers. Your 72 make a 15% share of voice. You appeared in 72 of 360 answers, so your mention rate is 20%. You don't score accuracy for competitors, because you can't check their records.
Add a notes column for the three most damaging wrong statements and the sources cited instead of you. Those notes become next month's fix list. And report the trend over three months before you call anything a win.
One last thing on the scorecard. Share of voice is an early signal. It isn't revenue. Wil Reynolds, who started Seer Interactive, shared a list of AI-era metrics in August that's a useful reminder of what to put next to it:
Here are the 11 metrics I wish I’d started tracking sooner, framed as questions: 1. Are branded organic searches growing faster than revenue?
That first question is a good one to copy. Put branded search volume from Search Console on the same sheet as your share of voice. Add a "how did you hear about us?" field to your enquiry form with ChatGPT and Google AI as options. If share of voice climbs and neither of those moves, something's off.
Spreadsheet or Tool: When Each Makes Sense
A spreadsheet is fine for a first baseline. A tool earns its keep once you need many prompts, several engines and repeat runs every month.
With 20 questions, four tools and three runs, you're looking at 240 answers a month. That's tedious by hand but doable, and doing it once yourself teaches you more about how these answers behave than any dashboard. My step-by-step AI visibility check is a good way to run that first pass.
Beyond that, tools make sense. Just be a demanding buyer. One SEO on r/bigseo was scathing about thin GEO platforms, and finished with advice I'd agree with:
If you're considering one of these platforms, just ask them to show you exactly what they do under the hood. Most of the time, you'll save yourself a lot of money and frustration.
That's one person's opinion, and many tools do careful work. But "under the hood" is the right test. Before you pay for any tool, ask four things. Can I see the exact prompts? Which engines and settings do you use? How many runs per prompt? How is each metric calculated? If the answer to the last one is "proprietary", keep asking.
The same questions apply if you're hiring someone to do the measuring for you. The GEO agency buyer's guide has the full list I'd take into that conversation.
Where ThriveFinity Fits
Share of voice is one of the measures in our AI search optimisation (GEO) work, alongside mention rate and accuracy. The question list, engines and number of runs are fixed before we start, and we publish how we handle repeat runs on our methodology page.
- AI Visibility Signal (free, under an hour): 12 buyer questions across GPT (OpenAI), Claude, Gemini and Perplexity, scored 0 to 100 and graded A to E. It's AI-only and unsigned, and a sensible first read before you build your own set.
- AI Visibility Diagnostic (£349, 3 working days): the measurement part of the Retrofit on its own, fully credited to the Retrofit if you go ahead within 30 days.
- AI Search Readiness Retrofit (£1,950, 10 working days): 100 to 150 buyer questions across five AI tools, your 10 most valuable pages fixed, and a re-test of the same questions at day 30. Signed off by a person.
- Quarterly Re-measurement (£595 a quarter): the same question set re-run each quarter, if you're not taking the monthly Retainer.
If you want to see what the output looks like first, the anonymised sample report shows share of voice by tool, the question map and the ranked fix list.
We're not the right choice if you need daily tracking across thousands of prompts. A dedicated monitoring platform is built for that. And if you only want the number once, the free Signal or a spreadsheet afternoon might be all you need.
❓ Common Questions
What is AI share of voice?
What is a good AI share of voice?
How many prompts and runs do I need?
Does Google Search Console show AI Overviews data?
How do you measure brand sentiment in AI answers?
Can ChatGPT do a sentiment analysis?
How accurate are AI Overviews?
Do AI Overviews hurt website traffic?
What tools track AI share of voice?
Sources
- Pew Research Center. Google users are less likely to click on links when an AI summary appears in the results. 22 July 2025 (browsing data from 900 US adults, March 2025).
- Search Engine Journal. Google AI Overviews now powered by Gemini 3. 27 January 2026, reporting Google's announcement.
- Hillel Maoz and Moshe Samet, Google Search Central. Introducing Search Generative AI performance reports in Search Console. 3 June 2026, updated for worldwide rollout on 31 August 2026.
- Microsoft Bing. Introducing AI Performance in Bing Webmaster Tools Public Preview. 10 February 2026.
- Google Search Central. AI features and your website. Checked 10 October 2026.
- Rand Fishkin, SparkToro. AIs are highly inconsistent when recommending brands or products. 27 January 2026.
- OpenAI Help Center. ChatGPT search. Checked 10 October 2026.
- Google Search Help. Find information in faster and easier ways with AI Overviews in Google Search. Checked 10 October 2026.
- ThriveFinity published prices and scope: /pricing (October 2026).