Everyone audits whether AI cites them. Almost no one audits what it is reading to decide.

/ 6 min read / By Faz / Updated July 28, 2026

There are two questions you can ask about your company and AI search, and they are not the same question. The first is whether the engine cites you, and how often, and how well. That is the output, and there is now a whole category of tools and a documented way to measure it. The second is what the engine is actually reading to build its picture of you in the first place. That is the input, and almost nobody audits it, which is strange, because the input is the part you can act on.

A visibility score tells you where you stand. It does not tell you why, and it does not hand you a list of things to fix. The source audit does both. When AI search says your product is the lightweight option, or files you in the wrong category, or quotes a price you have not charged in two years, that sentence came from somewhere specific. Find the somewhere, and vague reputation management turns into a short list of pages with URLs. This is how you build that list.

The thing that makes this tractable

Before the steps, the reason to bother. Most people avoid a source audit because they picture the entire web talking about them and no way to read all of it. The web is not talking about you that much. For a given buyer question, the engine reads a handful of pages, usually the same handful across runs, and its whole account of you traces back to a small set of sources: a review-platform category page, one or two editorial roundups, a community thread, maybe a news mention, and your own site. When you actually list them, it is closer to seven pages than seventy, and two of them are usually doing most of the talking. You are not boiling the ocean. You are finding the five to ten pages that write your answer.

One honest limit up front. You can only read the sources on the engines that show them, which means Perplexity, Google’s AI Overviews, and Copilot through Bing. Memory-based default ChatGPT will not show its work, so you cannot audit it directly. You infer its picture from the overlap in the engines you can see, because the sources that keep appearing across the visible engines are the ones most likely to have shaped the invisible one too.

Step 1: Run the buyer questions, not your name

Open the source-showing engines and run the real questions a buyer asks on the way to a decision. Not your brand name, which only tells you whether you exist. Run the category and comparison questions: best tool for this job, alternatives to the competitor you lose to, is your product good for this use case. For each answer, write down every source the engine cites. You are building a raw list of URLs, nothing more yet.

Step 2: Sort the sources into layers

Group what you collected. Review and directory platforms in one pile, editorial roundups and best-of lists in another, community and forum threads in a third, news and PR in a fourth, and your own pages in the last. Now look at which layer dominates your answers, because it varies by category and it changes the work. Some categories are decided on review platforms, some in communities the engine leans on, some on editorial roundups. The layer that keeps showing up is the layer you have to earn in, and it is often not the one you have been pouring effort into.

Step 3: Read the exact sentence, not the page

For each source that recurs, find the specific sentence the engine lifted. This is the move that turns an audit into a plan. The engine is not absorbing a whole page about you, it is carrying one line: the category it files you under, the adjective it repeats, the price it quotes, the weakness it names. Reading the origin sentence that wrote your label tells you exactly what each source is contributing to your answer, and whether that contribution is right, stale, or actively wrong. A source that describes you accurately is doing free work for you. A source carrying a two-year-old fact is the thing to fix, and now you know which page holds it.

Step 4: Score by leverage, not by how much it stings

Rank the sources, and rank them by leverage rather than by which one made you flinch. Leverage is three things multiplied: how often the source is cited, how wrong its sentence is, and how movable it is. An outdated comparison page on a site that takes correction requests is high leverage even if it is polite, because one email can change what it says. A heated community thread might sting more, but you control it less, so it sits lower. The instinct is to chase the loudest source. The audit exists to point you at the highest-leverage one instead, which is usually quieter.

Step 5: Write the source map

Turn the ranked list into one document: the pages that actually write your AI answer, in priority order, each with its origin sentence and what it should say instead. That map is the audit’s real output, and it is what every fix runs against. Whether the problem is a wrong category, a false fact, or a mislabel, the remediation is always aimed at specific sources, and this is the list of them. Rerun the audit on a schedule, because the source layer moves, and a page that was accurate last quarter can go stale without telling you.

What I got wrong

For a long time I called a citation tracker a source audit. I had a clean dashboard showing where clients were cited and how the numbers moved, and I mistook it for knowing what was going on. It measured the output beautifully and told me nothing about which pages to touch. The first time I actually sat down and listed the cited URLs for a client’s core buyer question, there were seven, and two of them were writing roughly eighty percent of the wrong story. Months of work would have been pointed at the wrong layer entirely if I had kept trusting the scoreboard.

The second thing I got wrong was leverage. Early on I let the loudest source set the agenda. A client had an angry, prominent forum thread and we spent real energy on it, while a dull, outdated comparison page that one email would have corrected kept feeding the actual mislabel across three engines. The thread felt urgent. The comparison page was the one moving the answer. Now the score is the first thing I calculate, and the loudest source rarely wins it.

Where this leaves you

The visibility number is worth having, but it is a symptom reading. The source audit is the diagnosis, and it is the difference between knowing you have a problem and knowing which five pages cause it. Run your real buyer questions on the engines that show their work, list the sources, read the origin sentence in each, score them by leverage, and write the map. Most of the reputation AI search reports about you is being written by a surprisingly short list of pages. The whole point of the audit is to stop guessing which ones.

If you want that audit run properly, the source map built and then worked page by page, that is the engagement, and it runs on a documented method.

Want this run for your B2B SaaS?

Founding pricing for the first 5 clients. Methodology fully public. Month-to-month, cancel anytime.

Apply to work together

Leave a comment

Your email address will not be published. Required fields are marked *