We ran an AI visibility report on our own brand: cited the most, recommended less

by Ross Gordon, Founder, Assist IQ

On 10 September 2026 we pointed an AI visibility report at one of our own businesses, the Scottish will-writing service ScottishWill. Twenty customer questions, each put live to four assistants: ChatGPT (GPT-5.6 Luna), Gemini (3.1 Pro), Claude (Opus 5) and Perplexity (Sonar Pro). One caveat on that list: the four models were chosen by the report tooling's own fixed rule for picking a web-search-capable model in each family, not by us, and they are not all the top tier of their families. Luna, for instance, is the entry tier of the GPT-5.6 family. Eighty cells, all eighty measured, none failed.

On the 60 unbranded answers that set the headline, counting each assistant's answer separately, the domain was cited as a source in roughly a third of them. The brand was named in under a fifth. That roughly two-to-one gap is the headline finding. The report also produces a composite score that weights mentions over citations; it is a scale specific to this report and not comparable to the six-pillar score on our own audit.

The assistants used our pages, then recommended somebody else

Across the full sample of 80 answers, the 20 branded ones included, scottishwill.co.uk was the single most-cited domain, ahead of the advice charities, a law firm, a bank and the other services that follow it. Counted by prompt rather than by answer, two thirds of the unbranded prompts had at least one assistant cite the domain, and a third had at least one assistant name the brand. Those figures come from our own run's data file, which we hold and have not published, so read them as one business's observation from a single run of 80 cells, not an audited industry figure.

The answers show the mechanism. The assistants pull the site's guide pages to explain how Scots law handles signing, witnessing and legal rights, use that to answer the legal half of the question, and then hand the buying half to someone else.

One control: all five branded prompts were answered correctly by all four assistants, so the entity is known once the name is in front of them. Reading the hedged answers, the pattern was consistent, and this is our reading rather than a field the report measures: an assistant holds back a recommendation when it cannot confirm, from what it can see, the facts it would want before naming a supplier. That is a facts problem on the pages rather than a ranking problem, and it is fixable.

What the report measures

We ran DataForSEO's Detailed AI Visibility Report skill, installed on 10 September 2026, driven from Claude Code. The vendor describes it as measuring "whether AI assistants actually name your brand when someone asks them what to buy", putting "a realistic set of customer questions to ChatGPT, Gemini, Claude and Perplexity" and recording "who gets named and who gets cited as a source". Running it needs a DataForSEO account with API credentials and the connection added to Claude Code once.

What we got out of it: mention rate and citation rate per assistant, a share-of-voice table against named competitors scored on the same answers, a ranked list of the domains the assistants cited most, the prompt-by-prompt grid showing which questions the brand is invisible on, and AI search volume per topic. The mention and citation numbers are reported separately, which is the split that matters.

What it cannot tell you

The report is honest about its own limits, and we added to the list while running it.

It is a snapshot. Assistant answers vary between runs, and with 60 unbranded cells one cell is worth 1.7 percentage points, so small movement against a later run is noise. This run is a baseline with no trend line behind it.

Claude cells need care. In the API's default mode, two Claude answers came back as quoted text with the source list stripped out. Re-calling both in raw mode changed the classification in both cases: one went from no citation to a confirmed citation of a real page. Every Claude call from the eleventh prompt onwards used raw mode, which means this run's Claude leg is itself mixed; the clean comparison starts with the next run, in raw mode from the first call.

One Perplexity answer was off topic. Asked which online will service to use in Glasgow, it answered about online banks. The call succeeded, so we counted it as a measured miss rather than re-rolling it, because re-running a successful call until it gives a nicer answer biases the sample.

And the tooling had a bug. The aggregation script that ships with the report skill normalised domains with a character-stripping call instead of a prefix strip, so any cited domain starting with "w" or a dot quietly lost its leading characters in one table. We patched our local copy of the script, corrected this run's data file, and rebuilt the report from the corrected data. If we had not opened the table, a client-facing report would have carried a misspelt domain.

The gap it exposed

Some prompts came back completely blank: no mention and no citation, on any of the four assistants. That is what the prompt-by-prompt grid is for, and it is a different problem from the one in the headline. Where the pages are already being cited, the job is to convert a citation into a recommendation. Where nothing appears at all, the job is to get into the answer set in the first place. The report's recommendations ran to six, three of them high priority, and they split across both.

Running this on a brand that is not ours

This report covers the buyer-question part of the six-pillar audit our AI visibility page describes, with a different mix of assistants. The way in is five questions that cost nothing first, and an AI Operations Day only later, if it makes sense. The day starts with a free fit call, where we "confirm the broad problem and whether a day inside the business is genuinely worthwhile". The AI Operations Day costs £750 per day on site. If you proceed with a qualifying pilot or build within 60 days, the full fee is credited against it.

Questions we get asked

What is the difference between a brand mention and a citation in AI search?

A mention is the assistant naming your business in the answer. A citation is the assistant using your page as a source, usually as a footnote or link. They move independently, and the gap between them is the useful number. In our own 10 September 2026 run, our domain was cited in roughly a third of the 60 unbranded answers and the brand was named in under a fifth of them, so the pages were trusted for the explanation while the buying recommendation went elsewhere.

What does an AI visibility report actually measure?

It puts a fixed set of realistic customer questions to several AI assistants, records which businesses each answer names and which pages it cites as sources, and reports the two rates separately per assistant. Alongside that it produces a share-of-voice comparison against named competitors scored on the same answers, a ranked list of the domains the assistants cited most, and a prompt-by-prompt grid showing exactly which questions the brand is invisible on. The grid is the part you act on.

Can one AI visibility report tell you whether your rankings have changed?

No. A single run is a point-in-time snapshot and assistant answers vary between runs. With 60 unbranded cells, one cell is worth 1.7 percentage points, so a few points of movement against a later run is noise rather than a trend. A first run is a baseline. It is only comparable to the next one if the settings match exactly. For us the clean baseline starts with the next run, because this one changed the Claude setting part-way through after we found the default mode intermittently dropped the source list.

More from the blog

We moved our AI operator to Claude Fable 5.1 the day after release and tracked it for a week

Anthropic released Claude Fable 5.1 on 1 September 2026. We switched our own AI operator onto it the next day and tracked seven days of usage: 160.8M tokens a day against an 89.3M baseline, and nothing throttled all week, on a counter we have never yet seen fire.

Read more

We scanned the AI skills our operator runs. Here is what a security scanner gets right and wrong

On 9 September 2026 we ran NVIDIA SkillSpector over the 58 skill directories in the two skill folders of our operator as they stood that day. Four scored CRITICAL. Nothing we read was malicious. Here is the routine, and the day we broke it.

Read more

Start with a free enquiry

Not a discovery call. Not a pitch with a calendar link. Five questions about how your business actually runs, and it costs nothing. Ross reads every enquiry himself and replies within one working day with a straight first answer: what looks worth automating, and what doesn’t.

If it looks like we can genuinely help, the next step is the AI Operations Day: one working day inside the business, followed by a written Opportunity Map. It shows what is hurting, what should stay human, and the best one or two jobs to prove first. That part comes later, and only if it makes sense for you.

Tell us what is slowing you down

Free to ask. No obligation. Ross replies personally within one working day.