Before You Buy an AI SEO Tool: 6 Questions to Ask the Vendor

In short: AI SEO tools fall into four categories — crawlers, visibility trackers, content optimisers and schema generators — and they answer four different questions. Buy the crawler first, because a visibility score of zero cannot tell you whether the problem is access or selection. Buy one tracker second, configured properly. Everything else can wait.
These products get marketed under several names — AI SEO software, AI visibility platforms, GEO tools, LLM visibility trackers, AI search monitoring — and the labels move faster than the capabilities do. What follows is organised by the job each one actually does, not by what the category is called this quarter.
Below: what each category can and cannot do, the six questions that force a vendor to show their working, which parts of the job no licence covers, and the order we would buy in if the budget were ours.
The Four Categories

Before comparing vendors, separate the market by the job each category performs.
| Category | What it does | What it tells you | What it cannot do |
| Crawler | Checks whether pages, links and directives can be reached and parsed by search and AI systems | Where access, indexing and technical barriers sit | Cannot decide what your brand should claim, and cannot guarantee AI visibility |
| Visibility tracker | Runs a defined prompt set and records appearances, citations and other visibility signals | How your site showed up for that specific prompt set | Cannot tell you with certainty why a result changed, and cannot represent every query a real user might ask |
| Content optimiser | Compares your content against selected topics, entities or competing pages | Where your coverage may be thin or duplicated | Cannot judge whether information is genuinely useful, accurate or strategically worth including |
| Schema generator | Produces structured data markup from the inputs you give it | Whether valid markup can be generated in a standard format | Cannot verify that an entity relationship is true, useful, or actually supported by the visible page |
Keep that distinction in mind when you buy, because each tool is answering a different question:
- A crawler asks: Can the system even reach and read this page?
- A tracker asks: What happened when we ran these specific prompts?
- An optimiser asks: What could we add or change in the content?
- A schema generator asks: Can this information be expressed in a machine-readable format?
Four useful questions — but four different jobs. Buying one doesn't answer the others.
The first one matters most, and Google's own documentation backs this up: passing the basic technical checks only makes a page eligible to be indexed. It gets you in line — it doesn't guarantee a seat. Google still decides — separately — whether to actually index the page and show it in search results.
Start With the Crawler, Not the Tracker
A tracker is tempting because its output is easy to understand. You enter prompts, receive a visibility score, and get a report showing where your brand appeared.
But a zero is not a diagnosis. If your content is missing from an AI-generated answer, at least two very different explanations are possible:
- The relevant content is accessible, but the system did not select it.
- The content was difficult or impossible for the relevant crawler or search system to access.
Those situations require completely different responses.
The first may lead to changes in content, entities, internal linking or source coverage. The second may require fixing robots.txt rules, access restrictions, rendering, status codes, indexing or other technical issues.
Google's current technical requirements are deliberately basic: Googlebot must not be blocked, the page needs to return a successful HTTP response, and it needs indexable content. Even then, indexing is not guaranteed. (Google for Developers)
That is why technical crawling should come before interpreting a visibility dashboard.

What about Google-Extended?
Google-Extended is a robots.txt control for how your content is used to improve Google's Gemini models and related generative AI services. Google's wording is that it limits "AI training and grounding in some of Google's other systems" (Google Search Central) — grounding here meaning the model retrieving your page at answer time rather than having learned from it in advance, though Google does not spell that distinction out. What Google does state plainly is that using Google-Extended "doesn't affect a site's inclusion in Search, nor do we use Google-Extended as a ranking signal in Search" (Google Crawling Infrastructure). It is not a switch for whether you appear in Google Search.
So a vendor should not present a Google-Extended setting as a direct control for AI Overviews visibility.
More broadly, robots.txt is an access-control mechanism for crawlers. It is important infrastructure, but it is not a universal "make my content appear in AI answers" setting. Google also provides other controls, including noindex, nosnippet, data-nosnippet and max-snippet, for managing how content can appear in Search and its AI features.(Google for Developers)
So what does control AI Overviews? There is a separate setting, and it is not in robots.txt. Google's Search generative AI control lives in Search Console under Settings, and it lets a site exclude its links and content from Search generative AI features — AI Overviews, AI Mode and the generative features in Discover. Google states that this control "only affects whether your content can appear in certain Search generative AI features" and "isn't used as a ranking or inclusion signal affecting other parts of Search". It is currently rolling out to a subset of website owners. That gives you three distinct levers, and vendors routinely blur them: Google-Extended (robots.txt) — AI training and grounding in Google's other systems. No effect on Search. nosnippet, data-nosnippet, max-snippet, noindex — how much of your page can be shown in Search and its AI features. Search generative AI control (Search Console) — whether you appear in Search's generative AI features at all. If a vendor offers to "manage your AI visibility settings", ask which of those three they mean. If they cannot name the right one, they are not the team to hand your access layer to.
Buying implication: if a vendor starts with a visibility score but cannot clearly explain what happens before that score is produced, ask for the crawling and access layer first.
What Visibility Trackers Are Really Sampling

The phrase "AI visibility" sounds broader than it really is. A tracker only measures the prompts it was asked to run. So the question that matters isn't whether the score is accurate — it's how the score was produced: who chose the prompts, how many times each one ran, and which language, region and engine they ran in.
A single run is an observation, not a finding. Results have to hold across repeated observations before they mean anything. Scores from different languages can't be compared directly either, because each language often draws on a different pool of underlying sources.
This isn't a translation issue. An English prompt and a Traditional Chinese prompt can return very different answers because they aren't retrieving the same documents. The sources pulled, the entities surfaced and the citations selected can all differ. That's why language and geography should be configured deliberately rather than left on default.
Before you compare two visibility scores, confirm the two methodologies match.
What Content Optimisers Get Right — and Where They Mislead

Content optimisation software has a legitimate use.
If several relevant pages consistently cover the same concepts and your page barely addresses them, that can be a useful editorial signal. Coverage analysis is especially helpful when a team is dealing with a large content library and needs a faster way to identify obvious gaps.
The problem starts when coverage becomes the goal.
Many optimisation platforms rely heavily on semantic overlap. The reason is simple: usefulness is hard to quantify, but overlap isn't.
If several competing pages repeatedly mention the same entities, the software often treats those entities as coverage requirements.
But repeated topics don't automatically create better content.
A page can mention every related concept and still be weak. Why? Because overlap is not the only thing that makes a page useful. A strong page may need to provide:
- a specific answer rather than a generic statement
- evidence for an important claim
- first-hand information
- clear definitions
- original examples
- meaningful comparisons
- context that competitors do not provide
A coverage score cannot automatically tell you which of those matters most.
Specific claims and vague claims are not interchangeable
Consider two sentences:
"Many businesses are investing more in AI search."
versus:
"A B2B SaaS company reduced its publishing cycle from six weeks to three after restructuring its approval workflow."
What sets them apart is evidence, not length. The second sentence is clearly more valuable — it offers a specific outcome that can be verified, not just a vague description.
A content optimiser may recognise topical overlap between the two — both mention entities and concepts tied to AI search investment or process improvement — but the commercial and editorial value of those statements is not equivalent. This is where teams should be careful about following automated recommendations at face value: a tool can confirm both pages "cover the topic," without telling you that one of them is doing far more work than the other.
The better question isn't:
"How do we get the score higher?"
It's:
"What information would make this page more useful, more defensible and more worth citing?"
That distinction matters because default metrics generally don't ask that second question for you. They can surface patterns worth investigating. They cannot replace editorial judgement about which of those patterns actually matter.
What Schema Generators Can't Decide

Schema generation looks deceptively simple. You provide information, the software produces JSON-LD, the validator reports no syntax errors, and the task appears finished. But valid syntax is the baseline, not the strategy.
Structured data is designed to help search engines understand page content and can make pages eligible for certain search features. Google recommends validating structured data and ensuring that it reflects information actually present on the page.
The harder question comes before the code:
What should you actually say?A generator can help express:
- who an organisation is;
- what a page is about;
- who authored an article;
- what a product is;
how entities relate to one another.
But it cannot independently establish whether a relationship is true. For example, software can generate an author relationship. It cannot decide whether the named person genuinely authored the article.
It can generate a product relationship. It cannot determine whether the product information shown in the markup is accurate and still matches the page. That judgement remains with the publisher.
Markup that does not match the page creates trust debt
Google specifically advises that structured data should match the visible content on the page. (Google for Developers)
That makes automated schema particularly risky when teams treat the generated output as something to paste everywhere.
The clean workflow is: decide the entity → verify the relationship → make the information visible → mark it up → validate it.
Not: generate everything → validate the syntax → publish.
A schema generator can save implementation time. It should not be mistaken for a semantic decision-maker.
The Three Jobs That Stay on Your Desk

Even with a strong technical crawler, a well-designed tracker and a useful content optimiser, three jobs remain fundamentally human.
1. Create something worth citing
Tools can identify what already exists. They are much less capable of deciding what your organisation should contribute that is genuinely useful.
That might be original research, first-hand experience, a proprietary dataset, a clear explanation of a difficult process, or a practical comparison built from real expertise.
If the underlying source is weak, better monitoring does not solve the problem.
2. Decide what you are actually claiming
Every page makes claims. Some are factual. Some are interpretive. Some are recommendations. Some involve numbers, products, people or organisations.
Someone needs to decide:
- What exactly are we saying?
- What supports it?
- How specific should the claim be?
- Is the evidence strong enough?
- Does the wording overstate what we know?
That is editorial responsibility, not software configuration.
3. Decide how far a rewrite should go
Identifying a gap is only the beginning. The harder question is deciding how to respond.
That could mean:
- one additional sentence;
- a new section;
- a rewritten explanation;
- a new source;
- a completely different page;
- or no change at all.
This is where budget expectations matter.
If a vendor implies that its platform can produce the source material, determine your strategic claims and make the final editorial decisions, ask what part of that promise is actually included in the licence. Those three jobs are usually still on your team.
Six Questions to Put to a Vendor

This is the part of the buying process that deserves the most attention. A polished demo can hide a lot of methodological differences, and these six questions are how you force a vendor to actually show their work.
1. Can we see the prompt list?
Start here. If a vendor won't show you the prompts behind their visibility score, there's no way to judge whether the sample makes sense. Ask for the actual wording, how prompts are categorized, the split between branded and non-branded queries, what language they're in, how often they get updated, and whether you're allowed to add your own. Without the underlying sample, a score is basically unauditable.
2. How many times is each data point run?
One generation is just an observation. It doesn't tell you much on its own. Find out whether the same prompt gets tested repeatedly, what happens when outputs differ from run to run, whether individual runs are stored or just averaged together, whether outliers get filtered out, and what happens to old scores when the methodology changes. If you don't understand how a number was made, comparing it to another number doesn't mean much.
3. Where is each prompt issued from?
"Global" isn't really an answer. Push for specifics: country, region, language, device, search environment, and whether any of that is configurable on your end. This matters a lot for international sites, since search behavior and available sources can look completely different from one market to the next.
4. Can you report each engine separately?
Once a vendor blends multiple AI engines into a single score, you lose the ability to see where a change actually came from — one engine, several, a specific prompt group, or a particular market. Ask for the breakdown. A single combined number is easy to read on a slide, but it doesn't tell you much when something shifts.
5. Can we export the raw records?
You want the ability to pull the underlying data — prompt, date, engine, location, response, citation or visibility outcome, whatever's available. Not because raw spreadsheets are exciting, but because your measurement shouldn't live entirely inside someone else's dashboard. It also makes it much easier to cross-check against Search Console, analytics, or your own editorial notes.
6. What happens to our baseline when your methodology changes?
Every tool evolves. That's fine. What's not fine is a vendor quietly changing their model, prompt set, or scoring formula and acting like the historical numbers are still comparable. Ask directly how they preserve your baseline through a methodology change. If the answer is just "the dashboard updates automatically," keep pushing — ask specifically what happens to your year-over-year comparisons. How they answer that tells you more about the vendor than the demo does.
A vendor who can answer all six is not necessarily the best product. But a vendor who cannot answer question 1 or question 6 is selling you a number you will never be able to audit or defend internally.
How the Tool Stack Maps to the Citation Loop

The Citation Loop is our own working name for the cycle we run — not a Google-defined framework and not an industry standard. On our AI SEO service page it has four stages:
Generative Visibility Audit → Knowledge Architecture → Schema & E-E-A-T Injection → Citation Monitoring
Software maps onto those four stages very unevenly, and that unevenness is the whole point of this article.
Generative Visibility Audit — this is where a crawler and a tracker both earn their licence fee. You need to know whether search systems can reach the site at all, and where the brand currently surfaces. Both jobs are measurement, and measurement is what software is genuinely good at.
Knowledge Architecture — partly software, mostly not. A crawler can show you orphaned pages and broken paths. Deciding which entities matter, which topics belong together and what the site should actually say is editorial work.
Schema & E-E-A-T Injection — a schema generator can produce the markup. It cannot decide whether the relationship in that markup is true, or whether the evidence behind an author or a claim is strong enough to be worth asserting.
Citation Monitoring — a tracker observes; it does not steer. It can show that a result changed after a content update. It cannot promise the next update will move the same way.
So two of the four stages are substantially served by software, and two are not. The half that isn't is strategy, editorial judgement and the quality of the underlying source — and no licence covers that.
What We'd Actually Buy, in Order

If the budget were ours, we would not start by buying every category. We would build the stack in this order.
1. Technical crawling
Start with the foundation. The priority is understanding whether important pages are accessible, indexable and technically sound. This is not glamorous, but it prevents the team from interpreting visibility problems that are actually access problems.
2. One correctly configured visibility tracker
Once the technical layer is in reasonable shape, add one tracker — not three. The goal is not to collect the largest possible number of dashboards.
It is to establish a repeatable measurement system with:
- a defensible prompt set;
- correct geography;
- correct language;
- appropriate run frequency;
- engine-level reporting;
- historical records.
The tracker becomes useful when the methodology is stable enough for changes to mean something.
3. Content optimisation, later and selectively
A content optimiser can be worthwhile once you have enough content and enough measurement to know where the editorial bottlenecks are. Use it to accelerate research and identify possible gaps. Do not make its coverage score the editorial KPI.
The best use is usually:
tool identifies a possibility → editor checks the evidence → writer decides whether the information belongs.
4. Skip the schema generator unless implementation is genuinely the bottleneck
Schema generation is one of the easier parts to handle manually or through existing CMS workflows when the underlying markup requirements are straightforward.
Google provides documentation and validation tools for supported structured data, and structured data still needs to reflect the visible page content. (Google for Developers)
If the organisation has a complex structured-data implementation at scale, automation may make sense.
But for a typical content operation, a dedicated generator should not outrank technical crawling or reliable visibility measurement in the budget.
Budget for the Part That Isn't the Licence
The recurring cost that's easiest to overlook is access-layer maintenance. Your robots.txt file doesn't become "done" forever, and neither does crawler access. Sites change, CMS platforms change, new subdirectories appear, development teams introduce staging rules, security systems change, and new AI crawlers and user-agent policies keep emerging. That's why crawler access needs to be treated as an ongoing operational task, not a one-time SEO setup.
Google describes robots.txt as a mechanism for managing how crawlers explore a site, and its current crawling documentation notes that Google-Extended operates as a separate control from Google Search inclusion. (Google for Developers)
For budgeting purposes, don't stop at licence plus implementation. The real number looks more like: licence + implementation + recurring access review + measurement governance — and that last part is usually where the actual internal cost sits. Someone has to own the prompt set. Someone has to review methodology changes. Someone has to investigate unexpected visibility swings. Someone has to check whether a technical change quietly altered crawler access. The software doesn't remove any of those responsibilities.
Buy the measurement you can explain, not the score that looks impressive in a demo.
FAQ
1. What should we buy first?
Buy technical crawling first. It tells you whether a visibility problem is actually an access problem, and a tracker cannot answer that. Add one correctly configured visibility tracker second. Content optimisers and schema generators come later, and only with a specific use case.
2. Can a tool replace an SEO agency?
No. Tools replace the repetitive parts: crawling, sampling, monitoring and content comparison. They do not replace strategy, original research, editorial judgement, technical implementation, or accountability for what your site claims. Ask which agency tasks are repetitive enough to automate.
3. Does Google-Extended control whether we appear in AI Overviews?
No. Google-Extended is a robots.txt control for AI training and grounding in Google's other systems, and Google states it does not affect Search inclusion or act as a ranking signal. The control that does affect AI Overviews is the Search generative AI control in Search Console.
4. Does a higher AI visibility score mean our SEO is better?
Not necessarily. A higher score means your site appeared more often for the specific prompts inside that tracker. It does not prove your content is more useful or your technical SEO is healthier. Ask how the score was produced before comparing it to anything.
Before You Buy Another Dashboard, Audit the Foundation First
At Maxlytics, we start by asking a different question: What exactly are you trying to measure?
Our AI SEO audits focus on the technical, measurement and content layers that influence AI search visibility. Before you commit to another platform, make sure you're solving the right problem.
資料來源:
- Google for Developers — Technical Requirements
- Google for Developers — Things to know about Google's web crawling
- Google Search Central — Intro to how structured data markup works
- Google Search Central — Optimizing your website for generative AI features on Google Search
- Search Console Help — Search generative AI control
- Google Search Central — AI features and your website

