AI Share of Voice: How to Measure Your Brand in AI Overviews and ChatGPT

AI share of voice is your slice of the brand mentions in AI answers, across a fixed set of buyer questions, against named competitors. Here's what to count, how many runs you need, what Google and Bing now report, and a monthly scorecard you can copy.

Pranav UnniFounder and lead verifier
Published
Updated
12 minRead time

In June 2026 Google added a new set of reports to Search Console. They show how often pages from your site appeared in AI Overviews, AI Mode and Discover's AI features. By 31 August they were live for every site.

It's useful data. But look at what it counts: your URLs, shown as sources. It doesn't tell you whether the answer named your company, what it said about you, or which competitors it recommended instead.

That gap is what AI share of voice fills.

I wrote this for marketing leads and founders at established B2B companies who've been asked "are we showing up in ChatGPT?" and want a number they can defend. It covers the five metrics worth tracking, what the new Google and Bing reports can and can't do, how many times you need to ask each question, and a scorecard you can copy into a spreadsheet this afternoon.

In short: AI share of voice is the share of brand mentions in AI-generated answers that belong to your brand, counted over a fixed set of buyer questions and compared with named competitors. Measure it the same way every month (same questions, same AI tools, several runs each) and read it alongside mention rate, recommendation share, citation share and accuracy.

What AI Share of Voice Means, and How It Differs From Search

AI share of voice counts who gets named in AI answers. In search, share of voice usually means your slice of visibility for a set of keywords, worked out from rankings and estimated clicks. An AI answer has no list of ten links to rank in. The assistant writes a paragraph and names a handful of companies, or none at all.

Google's AI Overviews are the AI-written summaries that appear above some Google results. Since January 2026 Gemini 3 has been the default model behind them, and people can ask a follow-up question and carry on into AI Mode. ChatGPT, Gemini, Perplexity and Claude write answers in a similar way, so the same measurement works across all of them.

Why bother? Because AI summaries change what people do next. Pew Research Center tracked the real browsing of 900 US adults in March 2025. About 18% of their Google searches produced an AI summary. When one appeared, people were more likely to stop browsing altogether.

Google visits where the user ended their browsing session
Google visits where the user ended their browsing session
Search page with an AI summary26%
Search page with only traditional results16%

Source: Pew Research Center, browsing data of 900 US adults, March 2025, published July 2025. US data. UK behaviour may differ. Chart by ThriveFinity.

If a buyer reads the answer and stops, being named in that answer is the visibility that counts. A page-one ranking they never scroll to doesn't help much.

One more distinction, because it trips people up. AI share of voice is a comparative number. It only means something next to the competitors you chose to track. If you pick weak competitors, you'll look great. Pick the companies you actually lose deals to. If you're not sure who they are, my guide to competitive intelligence for B2B companies shows how to find out.

The Five Numbers to Track

Five measures cover what most B2B companies need. Define them once, write the definitions down, and don't change them halfway through a year.

AI share of voice metrics
MetricFormulaWhat it tells you
Mention rateAnswers that name you ÷ all answersHow often you appear at all
AI share of voiceYour mentions ÷ mentions of all tracked brandsYour slice of the conversation against competitors
Recommendation shareAnswers that recommend you ÷ answers that recommend any brandWhether you're suggested or just listed
Citation shareCitations to your pages ÷ all citationsWhether AI tools use your site as a source
Accurate-mention rateAccurate mentions ÷ all your mentionsWhether what's said about you is right

Mention rate and share of voice can move in opposite directions, and that's worth understanding. Say a new competitor starts appearing in lots of answers. Your mention rate might hold steady while your share of voice drops, because the pie got bigger. Neither number is wrong. They answer different questions.

Recommendation share is the one I'd watch most closely for B2B. Being listed fifth in "companies that offer X" is nice. Being the one the answer suggests for "a 100-person firm like mine" is what puts you on a shortlist.

Citation share is the odd one out, because it measures your pages, not your name. An answer can cite your blog and recommend someone else. It can also recommend you while citing a directory. Track it, but don't confuse it with being chosen.

Semrush's explainer is a decent visual walk-through of the share of voice idea, and it ends with three ways to improve it, ordered from least to most effort. It's a vendor video built around their own toolkit, so treat the tool part as one option among several.

Video AI Share of Voice Explained: How to Track Your Brand in AI Models, Semrush on YouTube. Covers what the metric is, how to read it across different AI models, and three ways to improve it: topic gaps, third-party presence and sentiment.

What Google and Bing Now Report, and What They Don't

Both Google and Microsoft now give site owners some first-party AI data, and you should switch both on. Neither gives you share of voice.

Google's Search Generative AI performance reports launched on 3 June 2026. They show "how often URLs from your site appeared in generative AI features in Search and Discover", broken down by pages, countries, devices and dates. The announcement post was updated to say the reports reached all websites worldwide on 31 August 2026.

Microsoft got there first. In February 2026 it launched AI Performance in Bing Webmaster Tools, covering Copilot, AI summaries in Bing and "select partner integrations". It reports total citations, average cited pages, page-level citation activity and grounding queries, which are the phrases the AI used when it went looking for content to cite.

Aleyda Solis, the SEO consultant, summed up the Search Console version when she finally got access in July:

In the same post she notes it has no prompts, topics or clicks. Just impressions and URLs. Bing's version goes further, because grounding queries show you the sub-searches an AI ran before it cited you. This episode walks through what that data looks like and how to read it:

Video How to Use Bing’s AI Performance Data to Get More LLM Citations, Edward Sturm on YouTube. A breakdown of the AI Performance report in Bing Webmaster Tools, what grounding queries are, and how they differ from the prompts people type.

Here's how I'd use the platform data. It tells you which of your pages AI features pick up, and roughly how often. That's handy for spotting which pages to protect and which to improve. What it can't tell you is whether your company was named when a buyer asked "who should we hire for X?". A page can be cited in an answer that recommends a competitor. Your brand can be recommended in an answer that cites nobody's site at all.

So keep both. Platform reports for your pages. Your own prompt set for your brand.

Building a Prompt Set Buyers Actually Use

The prompt set decides everything that follows. A set built around your own brand name will flatter you. A set of vague industry questions will tell you nothing. Aim for 20 to 50 questions in four groups:

  • Category: "Best [category] providers for a [company type] in the UK".
  • Problem: the problem in the buyer's own words, before they know a category exists.
  • Comparison: "[You] vs [competitor]" and "alternatives to [competitor]".
  • Brand: "What does [you] do?", "Is [you] any good?", "How much does [you] cost?".

Take the wording from places buyers really speak: sales call notes, your Search Console queries, People Also Ask boxes and autocomplete. Guesswork from the marketing team tends to produce questions nobody types. If you want a starting list, the free AI Visibility Self-Audit worksheet has twelve buyer prompts you can adapt.

Keep brand questions to a minority. They inflate your score, because an assistant asked about you by name will nearly always mention you.

It also helps to know that a single question isn't really a single search. Google's documentation says AI Overviews and AI Mode may use a "query fan-out" technique, "issuing multiple related searches across subtopics and data sources" to build one response. Bing's grounding queries show the same thing from the other side. So when you write your set, include the sub-questions too: pricing, reviews, "for a company our size", "in the UK".

Once the set is agreed, freeze it. If you want to add questions later, add them as a new group and report it separately. Otherwise this quarter's number can't be compared with last quarter's, and the whole exercise loses its point.

Fix the settings as well as the questions. Write down, for each tool, whether you use the app or the API, whether you're logged in, which country and language you're in, and which model or mode is selected. Answers can differ by country, so test from where your buyers are. Change any of those settings and you've started a new series.

How Many Runs Is Enough?

More than one, and probably more than you'd like. AI answers change from run to run, so a single answer per question is an anecdote.

The clearest public evidence comes from SparkToro's January 2026 study. Six hundred volunteers ran 12 prompts through ChatGPT, Claude and Google's AI a combined 2,961 times. The same list of brands came back less than once in 100 runs. The same list in the same order was closer to once in 1,000.

The study also found something more useful. Some brands showed up again and again even while the order shuffled. One hospital appeared in 69 of 71 answers to its prompt. So the stable thing to measure is how often you appear across many runs, a rate, and the unstable thing is where you sat in any one answer. For accurate tracking, Rand Fishkin's advice in the write-up is blunt. You need to ask "over and over again", usually at least 60 to 100 times, and then average the results.

That's a consumer-scale recommendation, and most B2B companies won't run every prompt 60 times a month by hand. My advice is to pick a number you can sustain, run every prompt that many times in every tool, and never change it. Then treat small moves as noise until they hold for two or three months. A fixed method with modest runs beats a heroic one-off.

It helps to know how big a real change has to be before you believe it. Take the scorecard setup further down: 30 prompts, 4 tools, 3 runs, 360 answers. One extra answer naming you moves your mention rate by less than a third of a percentage point. A swing of five answers, which can easily happen by chance, moves it by about 1.4 points. So if your mention rate goes from 20% to 21.5% in a month, I wouldn't tell the board. If it goes from 20% to 30% and stays there, that's worth a conversation.

Rand explains the thinking behind this in a May 2026 interview, including why "AI rankings" are built on probability and what to track instead:

Video Why AI Rankings Don't Exist (And What To Track Instead) with Rand Fishkin CEO & Co-Founder SparkToro, Gen Furukawa @ SuperMarketers on YouTube. Rand on why AI answers are probabilistic, the two mechanisms that influence LLM visibility (retrieval and training data), and why PR and positioning matter more than tactics.

Measuring Mentions, Position and Citations

For every answer, you're recording three separate things: who was named, in what order, and which pages were cited. Here's the full method:

  1. Build a fixed prompt set. Write 20 to 50 questions real buyers ask, across category, problem, comparison and brand questions. Keep the list fixed so results can be compared month to month.
  2. Choose engines and settings. Pick the AI tools your buyers use, such as Google AI Overviews, ChatGPT, Gemini, Perplexity and Claude. Record country, logged-in state and app or API, and keep them the same each run.
  3. Run each prompt several times. AI answers vary between runs, so run each prompt more than once per engine and record every answer verbatim with its date.
  4. Extract mentions and citations. For each answer, list every brand named, its position, whether it was recommended, and which web pages were cited as sources.
  5. Score sentiment and accuracy. Score each mention of your brand positive, neutral or negative, and accurate, outdated or wrong, using a written rubric.
  6. Calculate and compare. Calculate mention rate, share of voice, recommendation share, citation share and accurate-mention rate for you and each named competitor, and compare against last month.

Position still matters within an answer, even if it's unstable across runs. Being named first in a list of five isn't the same as being fifth. For comparison prompts, also note whether the answer ends with a recommendation and for whom.

Record citations separately from mentions. Where an assistant searches the web, it usually shows its sources. OpenAI's help page says ChatGPT search responses "may include citations", and that ChatGPT "may search the web automatically when your question would benefit from current information." Those cited pages tell you which third-party sites you should be present on.

A caution, though. Rand Fishkin has argued that citation data can mislead you about why a brand gets recommended:

That's a view, not settled fact, and plenty of people find citation patterns useful. But it's a good reason to keep citation share and share of voice as separate columns. If the model already "knows" the brands it'll name, the work that moves share of voice is mostly about being widely and consistently described across the web, long before anyone types a prompt.

Brand Sentiment and Accuracy in AI Answers: How to Score Them

Score sentiment and accuracy as two separate columns, with a written rule for each label. Sentiment in AI answers is usually mild, so a five-point scale mostly adds noise. Three labels are enough:

  • Positive: the answer recommends you, or names a specific strength.
  • Neutral: you're listed or described without judgement.
  • Negative: the answer names a weakness or a risk, or steers the buyer elsewhere.

Then score accuracy on its own: accurate, outdated (an old price, product or location) or wrong. A positive but wrong mention is still a problem. The buyer will find out, usually on the first call.

Don't assume the answers are right just because they sound confident. Google's own help page for AI Overviews says plainly that they "can and will make mistakes", and labels AI responses as possibly including them. For your brand, the only fair test is to check each statement against your own records.

An AI model can do the first pass of labelling across hundreds of answers, and it'll save you hours. Check a sample by hand every month, keep the rubric fixed, and don't let the same tool both write the answers and mark them without a person looking.

Accuracy is also the metric you can influence most directly. Wrong answers usually trace back to unclear or inconsistent sources: an old pricing page that's still indexed, a stale directory profile, a description that differs from one site to the next. I go through how to fix those at the source in AI visibility for B2B brands.

A Monthly AI Visibility Scorecard (Template)

Use one row per brand and one tab per month. The figures below are illustrative, there to show the arithmetic. They aren't client data. The setup is 30 prompts, 4 AI tools and 3 runs each, which makes 360 answers.

Illustrative AI visibility scorecard, 360 answers
BrandAnswers naming brandMention rateShare of voiceRecommendedAccurate mentions
Your company7220%15%1858 of 72 (81%)
Competitor A19855%41%95n/a
Competitor B12635%26%40n/a
Others84n/a18%n/an/a

Here's how the numbers fit together. In total, 480 brand mentions were counted across the 360 answers. Your 72 make a 15% share of voice. You appeared in 72 of 360 answers, so your mention rate is 20%. You don't score accuracy for competitors, because you can't check their records.

Add a notes column for the three most damaging wrong statements and the sources cited instead of you. Those notes become next month's fix list. And report the trend over three months before you call anything a win.

One last thing on the scorecard. Share of voice is an early signal. It isn't revenue. Wil Reynolds, who started Seer Interactive, shared a list of AI-era metrics in August that's a useful reminder of what to put next to it:

That first question is a good one to copy. Put branded search volume from Search Console on the same sheet as your share of voice. Add a "how did you hear about us?" field to your enquiry form with ChatGPT and Google AI as options. If share of voice climbs and neither of those moves, something's off.

Spreadsheet or Tool: When Each Makes Sense

A spreadsheet is fine for a first baseline. A tool earns its keep once you need many prompts, several engines and repeat runs every month.

With 20 questions, four tools and three runs, you're looking at 240 answers a month. That's tedious by hand but doable, and doing it once yourself teaches you more about how these answers behave than any dashboard. My step-by-step AI visibility check is a good way to run that first pass.

Beyond that, tools make sense. Just be a demanding buyer. One SEO on r/bigseo was scathing about thin GEO platforms, and finished with advice I'd agree with:

That's one person's opinion, and many tools do careful work. But "under the hood" is the right test. Before you pay for any tool, ask four things. Can I see the exact prompts? Which engines and settings do you use? How many runs per prompt? How is each metric calculated? If the answer to the last one is "proprietary", keep asking.

The same questions apply if you're hiring someone to do the measuring for you. The GEO agency buyer's guide has the full list I'd take into that conversation.

Where ThriveFinity Fits

Share of voice is one of the measures in our AI search optimisation (GEO) work, alongside mention rate and accuracy. The question list, engines and number of runs are fixed before we start, and we publish how we handle repeat runs on our methodology page.

  • AI Visibility Signal (free, under an hour): 12 buyer questions across GPT (OpenAI), Claude, Gemini and Perplexity, scored 0 to 100 and graded A to E. It's AI-only and unsigned, and a sensible first read before you build your own set.
  • AI Visibility Diagnostic (£349, 3 working days): the measurement part of the Retrofit on its own, fully credited to the Retrofit if you go ahead within 30 days.
  • AI Search Readiness Retrofit (£1,950, 10 working days): 100 to 150 buyer questions across five AI tools, your 10 most valuable pages fixed, and a re-test of the same questions at day 30. Signed off by a person.
  • Quarterly Re-measurement (£595 a quarter): the same question set re-run each quarter, if you're not taking the monthly Retainer.

If you want to see what the output looks like first, the anonymised sample report shows share of voice by tool, the question map and the ranked fix list.

We're not the right choice if you need daily tracking across thousands of prompts. A dedicated monitoring platform is built for that. And if you only want the number once, the free Signal or a spreadsheet afternoon might be all you need.

❓ Common Questions

What is AI share of voice?
AI share of voice is the share of all brand mentions in AI-generated answers that belong to your brand, counted across a fixed set of buyer questions and compared with named competitors. Its simpler companion, mention rate, is the share of answers that name you at all. Track both, because a crowded answer can lift one and not the other.
What is a good AI share of voice?
There's no published benchmark worth trusting yet, and the number depends heavily on which questions and competitors you include. Compare yourself with the competitors you actually lose deals to, and with your own figure last quarter. A rising trend on a fixed question set means more than any single percentage.
How many prompts and runs do I need?
For most B2B companies, 20 to 50 buyer questions is a sensible start. Run each one several times in each AI tool, because answers change from run to run. SparkToro's January 2026 study suggests consumer-scale tracking needs dozens of runs per prompt. With a smaller budget, keep the setup fixed and report rates rather than single answers.
Does Google Search Console show AI Overviews data?
Partly. Since June 2026 Search Console has had generative AI performance reports, rolled out to all sites by 31 August 2026. They show impressions, pages, countries, devices and dates for appearances in AI Overviews, AI Mode and Discover. They don't tell you which brands were named in the answer, so they can't give you share of voice on their own.
How do you measure brand sentiment in AI answers?
Score every answer that mentions you as positive, neutral or negative, and separately as accurate, outdated or wrong. Use a written rubric so two people would score the same answer the same way. Then track the share of positive and accurate mentions over time instead of judging single answers.
Can ChatGPT do a sentiment analysis?
Yes. ChatGPT and similar models can label text as positive, neutral or negative, and they're handy for a first pass over hundreds of answers. Check a sample by hand, keep the rubric fixed, and don't let the same tool both produce the answers and grade them without a person checking.
How accurate are AI Overviews?
Google's own help page says AI Overviews "can and will make mistakes". Accuracy varies by topic. For your brand, the useful test is to run your buyer questions, record what AI Overviews say about you, and mark each statement accurate, outdated or wrong against your own records.
Do AI Overviews hurt website traffic?
They can reduce clicks. In Pew Research Center's study of 900 US adults (March 2025), people clicked a traditional result on 8% of visits to pages with an AI summary, against 15% without one, and ended their browsing session more often (26% against 16%). That makes being named inside the answer matter more.
What tools track AI share of voice?
Dedicated AI visibility tools run prompt sets across AI assistants on a schedule and report mentions, citations and sentiment. A spreadsheet works for a small set. Whatever you use, check that it shows you the prompts, the engines, how many runs it makes and how each metric is calculated.

Sources

  1. Pew Research Center. Google users are less likely to click on links when an AI summary appears in the results. 22 July 2025 (browsing data from 900 US adults, March 2025).
  2. Search Engine Journal. Google AI Overviews now powered by Gemini 3. 27 January 2026, reporting Google's announcement.
  3. Hillel Maoz and Moshe Samet, Google Search Central. Introducing Search Generative AI performance reports in Search Console. 3 June 2026, updated for worldwide rollout on 31 August 2026.
  4. Microsoft Bing. Introducing AI Performance in Bing Webmaster Tools Public Preview. 10 February 2026.
  5. Google Search Central. AI features and your website. Checked 10 October 2026.
  6. Rand Fishkin, SparkToro. AIs are highly inconsistent when recommending brands or products. 27 January 2026.
  7. OpenAI Help Center. ChatGPT search. Checked 10 October 2026.
  8. Google Search Help. Find information in faster and easier ways with AI Overviews in Google Search. Checked 10 October 2026.
  9. ThriveFinity published prices and scope: /pricing (October 2026).
Pranav Unni

Pranav Unni

Founder · ThriveFinity Connect on LinkedIn →

Pranav Unni is the founder and lead verifier of ThriveFinity. He reads and signs every paid deliverable personally, and writes about go-to-market decisions for established B2B companies.

Share this article

Free · under an hour

Get Your Starting Share of Voice

The free AI Visibility Signal runs 12 buyer questions across GPT (OpenAI), Claude, Gemini and Perplexity and scores how often you are named, 0 to 100.