Choosing an AI Visibility Tool: Brand-Mention Tracking vs Website Technical Scanning (AVS)
Table of Contents
- Separate the categories first: what are you actually buying
- One table
- Same name, different thing: there is more than one "AI Visibility Score"
- Three situations and which to choose
- One: the site has never had an AI-visibility technical check
- Two: the technical side is fixed and you need to report whether AI mentions you
- Three: an agency needs something repeatable to deliver to its own clients
- What the official documentation actually says
- Four common misunderstandings
- Summary
The 30-second answer: products sold as "AI visibility tools" fall into at least two categories. Do not compare their scores against each other.
(A) Brand-mention tracking: ask ChatGPT, Perplexity, Google AI Mode and others a fixed set of questions, and see whether the answers mention you, which page they cite, and your share against competitors. Semrush's July 2026 roundup defines these tools as ones that "track how your brand appears in AI-generated answers from ChatGPT, Google AI Mode, Perplexity, and similar platforms", and lists Semrush, Peec AI, Profound, HubSpot AEO and others.
(B) Website technical scanning: check whether your site can be opened and understood by search and AI retrieval: robots rules, structured data, llms.txt, citable content structure. The output is a score and a gap list. Our UltraProbe is in this category, as is Taiwan's TWTools three-axis check, which describes itself as 18 sub-metrics across SEO, GEO and AEO, rule-based scoring, no sign-up.
A tells you whether you get mentioned once you are in the room. B tells you whether the door is open and what to fix. A site that has never had a technical check should do B first, then decide whether to buy A.
Our position first: we are Ultra Lab. We maintain the open AI Visibility Score (AVS) and a free scanner, UltraProbe. We do B. This post cites our own guide and research as well as third-party and Google and OpenAI documentation, with a link every time; third-party product details are as of their sites today.
We have already written the pillar on what AI visibility is and how to measure it: the AI visibility guide. Score thresholds and method limits are in the AVS validation study. This post does not rerun experiments or re-teach structured data. It answers one thing: how to pick the category so you do not buy the wrong thing.
Separate the categories first: what are you actually buying
The same search term, "AI visibility", can mean entirely different products.
Mention tracking measures your brand's presence in answers. You give it fixed questions; it checks whether the model's reply includes you, your share against competitors, sometimes sentiment and cited URLs. Good for telling your manager "did AI mention us this week", but it usually does not tell you which line of robots.txt is blocking a crawler.
Technical scanning measures what can be checked automatically on the site. You paste a URL and get a score and gaps in seconds. Good for engineering and content owners planning fixes, but it does not tell you whether ChatGPT mentioned you today. That needs a fixed manual question set or a mention tracker, both covered in the visibility guide.
Ranking one against the other is like asking whether a blood-pressure monitor or a scale is more accurate. Both useful, different measurements.
One table
| (A) Brand-mention tracking | (B) Website technical scanning | |
|---|---|---|
| Question answered | Am I in the AI answer? How do I compare to competitors? | Can search and AI retrieval open and read my site? What is missing? |
| Typical output | Mentions, cited sources, share of presence (varies by product) | 0 to 100 score plus itemised gaps |
| Typical pricing | Mostly monthly subscription. Semrush's roundup lists its AI Visibility from $99 per domain per month and Peec AI's starter plan at $95 per month (July 2026; check each site for current pricing) | Free options exist; UltraProbe and TWTools need no sign-up |
| Cannot answer | Which line of HTML or robots to change | Whether a model mentioned your brand this week |
| When to use | Technical side is fixed, you want ongoing monitoring of answers | You do not know whether the door is open and need a fix list |
One note outside the table: UltraProbe also has a prompt-defence scan. That is a security category, separate from the visibility score. Do not merge them into one metric.
Same name, different thing: there is more than one "AI Visibility Score"
This is where people get caught.
Ultra Lab's AVS is a weighted score from an automated scan of your URL. Current v2.0:
AVS = SEO × 0.35 + AEO × 0.35 + AAO × 0.30
SEO covers classic search technical health, AEO covers answer-engine readability (structured data, llms.txt, AI crawler settings and so on), AAO covers accessibility to AI agents. This is the formula the live scanner actually computes.
The version matters. AVS started as a two-dimension v1 (SEO × 0.5 + AEO × 0.5), and the validation study used v1. So its findings, that roughly 60 percent of AI-cited sites scored grade B or above and that recommendation queries had a higher bar of about 80, were measured under v1; today's v2 scores will not line up exactly. We have added a version note to the study page.
Tracking tools have scores too. Peec AI's comparison page says Semrush's toolkit, in addition to showing aggregated mention counts, offers a separate 1 to 100 "AI Visibility score"; the page does not say how that score is computed, but it sits inside a mention-tracking toolkit and is not the same thing as a site-scanning AVS. HubSpot AEO calls its version a "brand visibility score". Peec itself measures visibility percentage, citation frequency and sentiment. In the same Semrush roundup, most tools revolve around mentions, citations, sentiment and competitors, but it also lists Botify (enterprise technical SEO) and Writesonic, which includes basic AI-readiness checks. Two categories filed under one name is exactly the point of this post.
So if someone says "our AI visibility is 78" and you say "our AVS is 78", those numbers are not comparable unless you first confirm both measure the same thing.
One more thing stated plainly: the relationship between score and being cited is correlation, not causation. Reaching a score does not mean AI will recommend you.
Three situations and which to choose
One: the site has never had an AI-visibility technical check
Do (B) first. Use UltraProbe or any rule-based scanner you trust to get the gap list, fix technical readability first by return on effort, then work on citable content. The order is in the visibility guide and the AEO guide; not repeated here.
The signal: you cannot say whether you have FAQPage structured data, whether you have llms.txt, or whether AI crawlers are allowed.
Two: the technical side is fixed and you need to report whether AI mentions you
(B) as the check-up, (A) or manual testing as the monitor.
On a tight budget: use the guide's manual method. Ten fixed questions, the same round of ChatGPT, Perplexity and Claude every month, watching the trend rather than any single result.
With budget, a large question set and a need for competitor comparison: then evaluate a mention-tracking subscription, and read the vendor's plan and data-source notes yourself.
The signal: GA4 already occasionally shows referrals from chatgpt.com, perplexity.ai or claude.ai, or manual testing has shown signal for several months. You need systematic tracking, not a first fix list.
Three: an agency needs something repeatable to deliver to its own clients
(B) is the easiest to turn into a repeatable deliverable: paste the URL, get the score and gaps, work through the list. Mention tracking usually carries separate licence and operating costs; check each tool's plans.
To describe the overall state of Taiwanese companies to a client, you can cite our Taiwan enterprise AI visibility report, but say it was measured at the time under v1, and do not present it as the client's score today.
What the official documentation actually says
Tool articles slide easily into "buy this and AI will recommend you". The official position is narrower. Align with it first.
Google's guide to optimising for generative AI features (updated 10 July 2026) is direct: its generative AI features "are rooted in our core Search ranking and quality systems", so SEO best practices still apply, and from Google Search's perspective optimising for generative AI search "is optimizing for the search experience, and thus still SEO". It also says many of the "hacks" suggested under AEO and GEO "aren't effective or supported by how Google Search actually works", and warns to be wary of third-party tools that promise ranking success or claim to use internal Google metrics.
On llms.txt, Google says Google Search itself does not use such files, and creating them "will neither harm nor help your site's visibility or rankings in Google Search"; it is completely fine to keep them for other services that use them. Structured data is similar: generative AI search does not require special markup, but it is a good idea to keep using it as part of overall SEO.
So when a technical scanner lists llms.txt as a check, the reason is that other AI services may read it, not that it earns points in Google.
OpenAI's crawler documentation lists four bots, each with its own job: OAI-SearchBot for ChatGPT search results, GPTBot for foundation-model training, ChatGPT-User for pages a user asks ChatGPT to visit, and OAI-AdsBot only for landing pages submitted as ads. The key sentence: sites opted out of OAI-SearchBot "will not be shown in ChatGPT search answers", though they can still appear as navigational links. Blocking GPTBot while allowing OAI-SearchBot is fine; the two settings are independent. OpenAI adds that for search results it can take about 24 hours after a robots.txt update for its systems to adjust.
This is exactly what a technical scan exists to catch: one line of robots configuration decides whether you exist in ChatGPT search, while a mention tracker will only tell you that you were not mentioned.
The conclusion, stated once: technical scanning handles "the door and readability"; mention tracking handles "your presence once inside". They complement each other, and neither alone is "AI marketing, done".
Four common misunderstandings
One: "Buy mention tracking and the site does not need fixing." Tracking can see "not mentioned"; it usually cannot see "because robots blocked this bot". If the door is closed, even an expensive tracker mostly reports a long blank.
Two: "A high AVS guarantees ChatGPT will recommend me." No. The validation study describes the score distribution of cited sites, a correlation, measured under v1. The score is a prioritisation tool, not a guarantee.
Three: "UltraProbe's score can be ranked against a commercial tool's AI visibility score." No. Same name, different thing; see above.
Four: "Having llms.txt guarantees appearing in AI answers." No. Google says its search does not read llms.txt, so having one or not does not affect Google ranking. Scanners list it because other AI services may read it and it is cheap. Do it if you like, but do not describe it as a guarantee.
Summary
Choose the category before the tool:
- You do not know whether the door is open and need a fix list: technical scanning (B), starting with UltraProbe.
- The door is open and you want to monitor your presence in answers: manual testing, or evaluate mention tracking (A).
- Do not rank the two kinds of score against each other. When reporting upward, keep "technical readiness" and "mention performance" separate.
This post comes from the team that maintains AVS and UltraProbe, as stated at the top. Third-party capabilities and prices are as of their own sites; our threshold figures are as published in our research, with the version noted, and are not a guarantee of recommendation.
Run a free scan first: UltraProbe. To understand the measurement method: the AI visibility guide. If you would rather not fix the gaps yourself, you can switch on UltraGrowth: the same gap list, as a monthly maintained subscription (setup from NT$19,800, from NT$6,800 a month; check the page for current pricing). Before deciding, read the SEO pricing and red-flags guide and compare deliverables item by item.
FAQ
What kinds of AI visibility tools are there?
At least two. The first is brand-mention tracking: ask ChatGPT, Perplexity and others a fixed set of questions and see whether answers mention you, which page they cite, and your share against competitors. The second is website technical scanning: paste a URL and check machine-checkable signals such as robots rules, structured data and llms.txt, producing a score and a gap list. They answer different questions and their scores cannot be compared.
Does a high AVS guarantee ChatGPT will recommend me?
No. Our validation study observed that sites cited by AI mostly sit in higher score bands, with a stricter bar for recommendation queries, but that is a correlation, measured under the v1 formula, not a causal guarantee. The score is for prioritising fixes, not a promise.
Is Ultra Lab's AVS the same as Semrush's AI Visibility score?
No, the names are just similar. Ultra Lab's AVS is a weighted score from a technical scan of your URL; the current v2.0 is SEO × 0.35 + AEO × 0.35 + AAO × 0.30. According to Peec AI's comparison page, Semrush's mention-tracking toolkit also offers a separate 1 to 100 AI Visibility score, which belongs to the category that tracks how a brand appears in AI answers. The two numbers cannot be compared.
Can I buy mention tracking and skip fixing the site?
You can, but if the site is technically closed to AI crawlers, a tracking tool will mostly keep telling you that you were not mentioned, without telling you which setting is blocking you. The safer order is to scan and open the door first, then decide on tracking.
Where can I get a free technical scan?
UltraProbe is free with no sign-up: paste a URL and get an AVS score with an itemised gap list at ultralab.tw/probe. Other rule-based scanners exist in the same category; pick the one whose report you find readable and that fits your workflow.