Say you ask ChatGPT who the best roofer in Warrington is, and it names you. A week later you ask again, and you’re gone.
Neither answer told you much. Each is a sample of one, and AI tools change their answers from one run to the next. Learning how to track AI search visibility means swapping the screenshot for a method: the same questions, asked the same way, often enough that the number means something.
If you’ve done the quick check in our AI search visibility guide, this is the next level. Eight steps take you to a fixed monthly routine in a free spreadsheet, the free reports Google and Microsoft now give you, AI visits in GA4, and a clear view on when a paid tracker earns its fee.
Step 1: Build a Fixed Set of Questions and Freeze It
Your question list is the ruler. Change it and every month-on-month comparison breaks, so build it once and lock it.
Aim for 15 to 20 questions, each with an ID (Q01, Q02 and so on). If you’ve done a GEO audit, start from that list rather than writing a new one.
Take the wording from customers, not yourself:
- Calls and enquiry emails. Note the exact words people use for the problem.
- Search Console. Check the queries that bring people to your service pages.
- Your contact form. Add a “What did you search for or ask?” field, an idea from Profound, one of the tracking vendors.
Group them so you can see which kind you win: service plus town, problem-led, comparison or “best”, and two or three branded prompts such as “What is [business] in Widnes known for?”. Ahrefs’ Louise Linehan recommends the same split of brand and category prompts.
For a pest control firm, an illustrative set might include “wasp nest removal St Helens”, “mice in my loft, who do I call in Warrington”, “best pest control company near Widnes” and “what do people say about [name]”.
Then freeze it. Add new questions as new IDs, never edit an old one, and review the set once a year.
You now have a numbered list you’ll ask, word for word, every month.
Step 2: Set Test Conditions You Can Repeat Exactly
Most home-made tracking is broken by the tester. Your history, memory and, above all, location change the answer, and it looks as if your visibility moved.
Write these conditions at the top of your sheet:
- Switch off personalisation. In ChatGPT, use Temporary Chat (OpenAI says it doesn’t use memory or save to your history) or log out. With memory on, ChatGPT can use saved details, including your location, to shape its searches. Ahrefs says the same: log out, use incognito.
- Fix your location. Google estimates where you are mainly from your IP address, and on phones also from GPS and Wi-Fi. Perplexity uses your network location if you haven’t set one. Test from your premises and you see what your neighbours see. Pick one place and one device, such as the office desktop on the office Wi-Fi. If customers are spread across Warrington and St Helens, name the town in the question rather than relying on “near me”.
- Use the same tools. ChatGPT (search on), Google AI Overviews and AI Mode, Gemini, Perplexity and Copilot. Note the model name if shown, because a model change is a legitimate reason for a jump.
- Test in the same week of each month.
A checklist on a sticky note beats a clever prompt.
Step 3: Run Enough Times for the Number to Mean Something
“Named in 4 of 10” is a start. But 10 runs can’t tell 30% from 50%.
If you appear in half of 10 runs, your true rate could sensibly be anywhere from about 20% to 80%. The range narrows as runs add up:
| Runs | Rough margin of error |
|---|---|
| 10 | about ±31 points |
| 30 | about ±18 points |
| 60 | about ±13 points |
| 100 | about ±10 points |
These are rough guides from the standard margin-of-error formula, worked out at a 50% hit rate, not findings from a study. The formula gets loose at small samples, so read the 10-run row as a warning, not a precise range.
SparkToro’s researchers suggest 60 to 100 runs before the data means much, and found narrow categories give tighter results than broad ones. A roofer in one town is a narrow category, which helps you.
You don’t need 100 runs of every question. Pool them. Twenty questions, three runs each, gives 60 answers per tool per month: enough for a per-tool share you can compare month to month. Per-question figures stay rough, so use them to spot patterns, not to celebrate.
Five tools at 60 answers each is 300 answers a month by hand. If that’s too much, cut tools before you cut runs.
Fewer tools, more runs.
Step 4: Record Every Answer in One Simple Sheet
A folder of screenshots can’t be counted. One row per answer can. This is the core of free AI visibility tracking, and it works in Google Sheets or Excel.
On a tab called Raw, set up these columns:
Month | Date | Tool | Question ID | Run no. | Named? (Y/N) | Linked? (Y/N) | Page linked | Description accurate? (Y/N/Partly) | What it said | Competitors named
Keep three things separate, because they fail in different ways (the AI search citations guide covers mentions versus citations):
- Named is the recommendation itself.
- Linked is the only one that can send a visit, and shows which of your pages the tool used.
- Description catches the answer that names you and gets it wrong: the wrong town, or a service you’ve dropped. Ahrefs’ checklist records accuracy and tone for this reason.
On a Summary tab, your share of ChatGPT runs where you were named in October is:
=COUNTIFS(Tool,"ChatGPT",Month,"2026-10",Named,"Y")/COUNTIFS(Tool,"ChatGPT",Month,"2026-10")
Here Tool, Month and Named are named ranges for those columns. In plain terms: answers that named you, divided by all answers, for that tool and month. Copy it for each tool, then repeat with the Linked column. A pivot table of the Competitors column shows who AI tools actually name alongside you.
Screenshot anything wrong about you. That’s your evidence when you go to fix the source.
You now have one tab of raw rows and a summary with a share per tool per month. If you keep only one column beyond Named, keep Description.
Step 5: Read the Trend, Not the Month
ChatGPT’s share drops from 40% to 30% and it feels like a problem. It probably isn’t.
Use two rules:
- Ignore small moves. At about 60 answers per tool a month, treat a change of under roughly 15 points as noise. That’s the Step 3 margin, rounded up: a rule of thumb, not a statistical test.
- Watch three months. A rolling three-month average smooths the bounce. Three months moving the same way is a trend, even if each step is small.
Check the competitors column before reacting. If everyone fell in the same month, the tool probably changed, not you. That’s why Step 2 has you noting the model name.
Don’t turn any of this into a position. SparkToro’s Rand Fishkin is blunt that a “ranking position in AI” isn’t a real thing, because the answers change too much from run to run to rank in.
Descriptions are the exception to the patience rule. A new wrong fact is worth chasing the month it appears.
Google rank tracking asks where you are. This asks how often you’re there, and whether that’s rising.
Step 6: Add the Free Data Google and Microsoft Now Give You
Google and Microsoft now report AI visibility from real searches, free.
Search Console’s Generative AI performance report counts impressions, which Google defines as “how many times links to your site were shown to a user in a generative AI feature on Google Search”. It covers AI Overviews and AI Mode, by page, country, device and date, with no clicks, click-through rate or queries. Testing began with UK site owners on 3 June 2026, and it reached all sites worldwide by 31 August 2026. Sites without enough AI impressions may not see it yet.
Bing Webmaster Tools AI Performance, in public preview since 10 February 2026, counts citations of your pages across Microsoft Copilot, AI summaries in Bing and some partners, plus the grounding queries it searched with. No clicks.
Both count links to your pages, not mentions of your name, and cover only their own tools. They add to your test runs; they don’t replace them.
Ten minutes a month, two free numbers.
Step 7: Connect AI Visibility to Enquiries in GA4
Visibility only matters if it turns into calls. You may have heard AI visits vanish into “direct”. Some do, but GA4 now separates some out.
In May 2026 Google added an AI Assistant channel to GA4’s default channel group. Google’s definition covers visits “from sources like ChatGPT, Gemini, Deepseek, Copilot, or Grok” and says it “excludes Google’s AI Overviews and AI Mode.” Perplexity isn’t named.
A custom channel group catches more:
- Go to Admin > Data display > Channel groups and select Create new channel group.
- Add a channel called AI where Source matches a regex such as
chatgpt\.com|perplexity\.ai|gemini\.google\.com|copilot\.microsoft\.com|claude\.ai. - Move it above Referral. GA4 assigns traffic to the first channel it matches.
- Save. Custom groups apply to past data too, and a standard property can have two.
ChatGPT links usually arrive tagged utm_source=chatgpt.com.
Some visits still hide. AI Overviews and AI Mode clicks sit inside Organic Search, and links opened from apps or messengers, or copied and pasted, land in Direct. Seer Interactive calls the AI Assistant channel “a floor on your AI traffic, not a ceiling.”
Now filter your key events (calls, form fills) by the AI channel. Our local SEO tracking guide covers setting those up. Add “AI assistant (ChatGPT etc.)” to your “How did you hear about us?” options too.
Record AI sessions and AI enquiries as two more columns on your summary tab.
Step 8: Decide Whether a Paid Tracker Is Worth It
Paid AI search monitoring tools automate Steps 3 and 4. Are they better than your spreadsheet?
Pay for one if you have several locations, a question set too big to run by hand, or three months of manual data you now want automated. Not before you’ve frozen a question set.
It’s a young market. Trackers include Profound, Peec AI, Otterly.AI and Trackerly; SEO suites offer Semrush’s AI Visibility Toolkit, Ahrefs Brand Radar and SE Ranking. Ask any of them:
- API or the real interface? Surfer’s August 2026 test of 1,000 prompts (one sample each) found brands named via the API and in the consumer interface overlapped only about 16% to 24%, and ChatGPT’s interface showed about four times as many sources. Your customer sees the interface.
- How many runs per prompt, and will you show the maths? Fishkin says to make sure your provider “shows their math”.
- Can I set the location to my town, not just “UK”?
- Will it match my business name reliably? One tracker’s developer found exact-name matching breaks on small local businesses.
Any tool reporting a single “AI rank” fails question 2 by definition.
If you’d rather have this run for you each month, that’s part of what we do. Tell us what you do and where, and we’ll give you a straight answer on whether we can help. If we can’t, we’ll say so.
Best for multi-site or time-poor owners. Skip it if you haven’t frozen a question set yet.
FAQ
How often should I track AI search visibility?
Monthly, in the same week, under the same conditions. Weekly checks at small sample sizes mostly measure noise. You can glance at the free Search Console and Bing Webmaster Tools reports more often, because they count real searches rather than your own test runs.
Can I track brand mentions in ChatGPT for free?
Yes. Use Temporary Chat or log out, ask a frozen question set several times, and log one row per answer in a spreadsheet. GA4’s AI Assistant channel then shows the visits ChatGPT sends, and your key events show which became enquiries.
Why does a tracking tool show something different from what I see?
Different location, account state, run or collection method. Tools that query the API can see very different answers from the app customers use, as Surfer’s comparison found. Neither is wrong on its own, so compare like with like.
Is an AI visibility score meaningful?
Only if you know what it’s made of: which questions, how many runs, which tools and which location. A share of repeated runs across a fixed question set is meaningful. A single score or position with no method behind it isn’t.