The question usually arrives as a shopping question. Someone on the team has three tabs open comparing AI visibility platforms, there is a budget line waiting to be spent, and the ask is “which one should we buy?”
It is the wrong first question. Not because the tools are bad, but because buying one is the most natural way to feel like you have started without actually starting. The honest version of the question is a sequencing question: what does a tool do, what does it not do, and at what point in the work does it start paying for itself?
Here is the answer I give clients, including the ones who arrive having already bought one.
What a tool actually does
Every AI visibility platform on the market, whatever it says on the pricing page, does one core job: it runs queries through AI engines on a schedule and records what comes back. Who got named. Who got cited. How you were described. How that changes week over week.
That is measurement. It is genuinely valuable, for a reason that surprises people the first time they check by hand: AI answers are not stable. Ask the same engine the same question three times and you can get three different answers. A single manual check is a coin flip, which is why measuring AI search visibility properly means a fixed query set, repeated runs, and multiple engines. Doing that by hand for five queries is a morning. Doing it for fifty queries across five engines every week is a job. Tools exist because that job is real.
So the tool question has a real answer. It is just narrower than the sales page implies.
What a tool cannot do
Now list what actually moves your visibility in AI answers. Not what measures it. What moves it.
The engine describes you based on what credible third parties publish about you. So the levers are things like: a comparison page on someone else’s site that still carries your old positioning, and the outreach email that gets it corrected. A review profile that has gone stale. A category page that files you under the wrong noun. Content on your own site written specifically to be liftable and citable, which is a planning and writing discipline, not a setting. The slow work of earning mentions you do not control.
None of that happens inside a dashboard. A tool can tell you that the engine calls you “the lightweight option.” It cannot find the three pages that taught the engine that phrase, judge which one the engine actually leans on, write the correction, or build the relationship that gets the correction published. The finding is automatable. The fix is not.
This is the honest split: tools measure, content and source work move. A tool with nothing behind it is a scoreboard for a game you have not started playing.
The trap is the feeling of completion
I have watched this pattern enough times to call it a pattern. A team buys a visibility platform, spends two weeks setting up query sets and dashboards, presents the baseline report internally, and feels the satisfying click of a project shipped. Then the dashboard sits there, faithfully recording a number that nothing is acting on, and three months later someone asks why the number has not moved.
It is the same trap I wrote about with llms.txt: a legible, finishable task stands in for the unglamorous work, and completing it feels like doing AI search optimization. Buying the scale is not the diet. The tool purchase is the single easiest item on the whole program to complete, which is exactly why it so often gets completed first and alone.
There is a second, quieter cost. Tool budgets and work budgets usually come from the same pocket. The team that spends a meaningful monthly sum on monitoring has, in my experience, a harder time finding budget for the writing and outreach the monitoring is supposed to measure. You end up precisely instrumented and structurally idle.
The line where a tool starts paying for itself
There is a real crossover point, and you can locate it with two questions.
First: do you have anything worth measuring yet? If you have a handful of pages, no query map, and no source work in motion, a weekly tracking report will tell you the same thing every week: not much, everywhere. You do not need software to know the score is zero. Hand checks are not just cheaper at this stage, they are better, because typing the queries yourself and reading the full answers teaches you how your buyers’ questions actually get answered, which no rollup metric preserves.
Second: has manual checking become the bottleneck? Once there is a real query set built from a buyer query map, content shipping against it, and fixes in flight that need before-and-after evidence, the measurement job grows past a spreadsheet. When someone on the team is spending hours a week running the same prompts, or when you skip a scheduled check because nobody had time, the tool is now cheaper than the labor it replaces. That is the buy signal. It is a capacity signal, not a maturity badge.
For most B2B SaaS teams I have worked with, that crossover arrives somewhere in the first few months of a real program, not on day one. And when it arrives, the useful tools are the measurement ones. I keep a current list of the tools I actually use for client work, and the honest summary of that list is: a few things for tracking, and a lot of ordinary writing, research, and email.
What I got wrong
Early on I recommended a client buy a monitoring platform in week one of an engagement. My reasoning sounded rigorous: we need a baseline, we need trend data, we should instrument before we intervene. All true in principle.
In practice the dashboard spent its first two months reporting a flat line at zero, because it takes time for new content and corrected sources to show up in answers, and we had only just started making them. The baseline could have been established with one afternoon of hand checks. Meanwhile the subscription cost, multiplied over those months, was roughly the cost of the two source corrections we had deferred. We had bought precision for a period when nothing needed precise measurement, with money that the moving parts needed.
Now I sequence it the other way. Hand checks to baseline. Content and source work first. The tool when, and only when, the checking becomes the constraint.
The decision rule
If you are choosing between spending on a tool and spending on the work, spend on the work. The question “tool or content?” almost answers itself once it is phrased honestly: one of them can exist without the other and still produce results, and it is not the tool.
If you are further along, shipping consistently, fixing sources, and straining to measure it all by hand, buy the measurement and keep the work running. Both answers are cheap to check against your own situation: look at last month. If more hours went into looking at dashboards than into making and fixing the things the dashboards watch, the sequencing is backwards.
This decision, tools versus content versus source work and in what order, is most of what I do for clients. The full sequence is public on the methodology page, and if you would rather have it run for you than run it yourself, that is what the retainer is for.