Methodology

The CITE Method, documented in full.

Four steps that get a B2B SaaS company named, described correctly, and cited in the AI answers its buyers read. Documented in full so you can run it yourself or hire me to run it.

Most B2B content gets published to fill a calendar, not to win customers. The companies winning in 2026 publish less and think harder. Six well-built pieces beat thirty generic ones, every time. CITE is the framework for changing that answer: Capture buyer questions, Investigate who owns the answers today, Target what to build and how, Earn citations in the sources you do not own.

This page is the actual methodology. Not a teaser. If you read it carefully and have a strong content team, you can implement it yourself. Plenty of clients read it carefully and still hire me, because reading the playbook and running it consistently for six months are different problems.

Most agencies hide their methodology. The methodology is the product. Hiding it would mean there isn't much there.

C  /  Capture buyer questions

Every program starts with a map of what your buyers actually ask, across Google, ChatGPT, Perplexity, Reddit, and the communities they hang out in. Most agencies start with keyword volumes. I start with questions. Keywords describe what people search. Questions describe what people want to know. The difference shows up in every piece that follows.

The question harvest pulls from sources in priority order:

  • Customer and sales call recordings (the richest source of real buyer language)
  • Top question-format posts in the primary subreddits and forums
  • Quora questions with real engagement
  • Competitor blog comment sections and review platforms
  • G2 review fields where buyers describe what they wish was different
  • Direct prompts to ChatGPT, Claude, and Perplexity to surface the questions those engines already answer

Candidates are scored across five dimensions:

  1. Buyer intent. Would someone asking this be ready to buy in 0 - 90 days?
  2. AI prevalence. Does ChatGPT or Perplexity actually return a substantive answer? Some questions sound great but engines just return "ask a professional."
  3. Movement potential. Is the current top answer beatable? Are cited sources weak (random blogs) or strong (top G2 listings, viral Reddit threads)?
  4. Strategic value. Does ranking here actually affect the client's business?
  5. Coverage gap. Does the client currently not appear?

The output is the buyer query map: a documented, prioritized list of every question worth answering, scored, mapped to the surface where it appears, and split into the questions that get full flagship treatment vs. the questions that get shorter pieces.

If you want to see the buyer query map step in action, the public DIY version is documented step by step in how to run your own AI search visibility audit, with twenty prompt patterns and a free Google Sheet template.

Why this matters

Most content programs skip this step. They guess what buyers care about and write toward keyword volumes. Six months later, nothing ranks and nothing gets cited. Capturing real questions first is the single biggest difference between content that earns its place and content that fills a calendar.

I  /  Investigate who owns the answers

Where do your competitors show up in Google and AI search? What do they have that you don't? Where are the gaps you can win? This produces the priority list. The work isn't beating competitors at what they're already winning. It's finding the questions and surfaces they've ignored, and getting there first.

For each question on the buyer query map, I run the prompt three times on each of four engines: ChatGPT (with search enabled, logged out, US locale), Claude (with web search), Perplexity (default mode), and Google AI Overviews (clean browser, location matched to the buyer geography).

Why three runs

LLM responses are non-deterministic. One run is noise. Three runs averaged out is a real signal. This is the single most important methodological discipline. Agencies that report from a single run are reporting from a coin flip.

For each run I capture: did the client appear, where in the answer, what URL the engine cited, the sentiment, and which competitors showed up alongside. The output is the competitor gap report: a documented map of where competitors appear, where you don't, which sources the engines actually cite, and where the recoverable gaps are.

Every run is scored on the same five-tier scale, because "did we show up" is not a useful answer:

  1. Named first. The engine recommends you ahead of everyone else, usually with a reason attached.
  2. Named among several. You are on the shortlist, in no particular order.
  3. Named but qualified. You appear with a condition attached. "Good if you are small." "Fine for simpler use cases." Present, and quietly losing.
  4. Cited but not named. The engine used your page as a source and recommended somebody else. You did the work and a competitor got the sentence.
  5. Absent. Not in the answer at all.

Tiers three and four are the ones teams never find on their own, because a dashboard that counts mentions reports both of them as a win.

Alongside the score, every run records the exact URLs the engine cited. Those URLs are the real output of this step. Aggregated across every priority question they produce the source map: the ranked list of pages currently writing your category's answer. In most categories it is a short list. A handful of domains usually account for the bulk of citations on the questions that matter, and that list becomes the target for the last step.

T  /  Target what to build, and how

Six pieces a month, split between two flagship long-forms (the ones that get recommended by ChatGPT and earn links) and four shorter pieces (the ones that rank in Google for specific buyer searches). Both compound. Neither dominates. This balance is what makes a six-piece program outperform a thirty-piece one.

Flagship pieces (2,500+ words each, two per month):

  • Built around the highest-priority buyer questions from the query map
  • Original takes, original data, deep research
  • Designed to be the best piece on the topic, not a competitor copy
  • Earn inbound links and get recommended by ChatGPT and Perplexity

Shorter pieces (1,200 - 1,800 words each, four per month):

  • Built to rank in Google for specific commercial searches
  • Capture buyers actively comparing options or troubleshooting
  • Tighter scope, faster turnaround, structured for AI extraction

The output is a monthly content plan, refreshed each cycle. Plans don't get locked in for 12 weeks. Each month's plan responds to what moved in the prior month: which questions got citations, which surfaces produced inbound, which formats outperformed. The discipline is the framework. The plan adjusts.

Engineering each piece so an engine can lift it

Each piece is built to rank in Google, get recommended by ChatGPT, and work as social content. Not three different pieces. One piece that performs in three places. This is the production discipline that justifies the price. Generic content can rank or get recommended by accident. CITE-built content does both on purpose.

What goes into engineering each piece:

  • The answer in the first 100 words. AI engines recommend answers, not setups.
  • Headers phrased as questions where natural ("What is X?", "How does X work?", "When should you use X?")
  • Original data, surveys, or your own frameworks. AI heavily prefers original sources over recycled ones.
  • Real attribution and credentials visible. Bylines with named operators get recommended more than anonymous posts.
  • Schema markup where applicable: Article, FAQ, HowTo. Recommendations correlate with structured data.
  • Named tools, platforms, and people the AI already knows. Co-mentions build the entity graph that includes the client.
  • Internal linking that connects flagship pieces to the shorter pieces, building topical authority.

A flagship that gets cited will outperform a polished piece that doesn't. Every time.

E  /  Earn citations you do not own

This is the step most programs skip, and it is the one that moves answers. Everything up to here happens on your own domain, and your own domain is the weakest source an AI engine has. A company describing itself is the least trusted input in the pool.

The answer your buyer reads is assembled from third-party pages: comparison articles, review profiles, practitioner write-ups, community threads, regional publications. If those pages do not mention you, no amount of editing your own site puts you in the answer. You are not ranked low. You are absent.

So the final step is a source-layer operation, run every month:

  • Identify the exact pages the engines cite for your priority questions. In most categories the answer traces back to two or three sources, not twenty.
  • Correct what is wrong at the origin. Stale comparison pages, an old price, a feature you shipped last year that a reviewer still lists as missing.
  • Earn mentions in the sources the engines already read, through manual outreach to the publications, roundups and comparison pages that own your category answer.
  • Publish the proof those sources need in order to cite you: original data, a documented method, a result someone can check.
  • Show up honestly in the communities your buyers read, as a named person with a history, not an account that appeared last week.

None of this can be faked. Coordinated mentions get detected, and a busted attempt becomes a cited source working against you. It is slower than publishing, harder to sell, and it decides whether the other three steps show up in an answer at all.

You cannot fix your reputation on your own website. That is the entire reason this step exists.

The four moves, ranked by leverage

Not by how much control you have over them. Control and leverage point in opposite directions here, which is why most programs spend their hours in the wrong place.

  1. Correct it at the origin. A stale comparison page or an out-of-date review profile that the engines keep citing. One email to the author, with the correct fact and a reason to care, can change an answer that no amount of publishing would have touched. Highest leverage, lowest control, and the move almost nobody makes.
  2. Earn corroboration in sources the engine already trusts. The publications, roundups and comparison pages that show up in your source map. Real outreach, on the merits, one at a time.
  3. Publish the proof those sources need. Original data, a documented method, a result that can be checked. This is the on-domain work, and its job is to give the off-domain sources something worth citing.
  4. Own the comparison query. The alternatives and versus pages, written honestly enough that an engine treats them as a neutral reference rather than marketing.

What a month of this actually looks like

Unglamorous. The source map from the previous step gives a target list, usually short. Each target gets read properly, so the outreach references the actual page and the actual error rather than a template. Most of it is email. Some of it is a correction request. Some of it is offering a data point the page is missing.

Reply rates are what reply rates are. A month where two pages get corrected and one new source mentions you is a good month, and those three changes can move a category answer more than six pieces of content will. Nobody can promise a number here, which is exactly why it is easier to sell content by the piece.

What I will not do

  • Paid placements dressed up as editorial coverage.
  • Seeded community threads, sock puppets, or any coordinated mention campaign. Engines detect these, and a busted thread becomes a cited source working against you permanently.
  • Review-site manipulation of any kind.
  • Promising a specific citation, on a specific engine, by a specific date.

The honest version of this step is slow and partly outside anyone's control. That is a real limitation of the work, not a sales problem to write around.

What the first 90 days look like

The four steps describe the method. This is the calendar they run on, so you know what you are buying in week two rather than month four.

Week 1. ICP and questions. We agree who the buyers actually are before anything else, because every later decision is made against that map. Sales call recordings, existing customer language, the segments you win and the ones you lose. Then the question harvest starts.

Week 2. Baseline and source map. The priority questions get run three times each across four engines. You receive the buyer query map and the competitor gap report, including the five-tier scorecard and the ranked list of pages currently writing your category answer. This is the point where most teams see the problem clearly for the first time.

Weeks 3 and 4. First publishing and first outreach. The first pieces ship against the highest-priority questions, and outreach begins on the two or three sources at the top of the map. Both start in month one. The off-domain work is not phase two.

Month 2. The flat stretch. Six more pieces, continued outreach, and the first re-measurement against the baseline. Expect the headline number to look unchanged. This is normal and it is the part worth naming in advance: retrieval engines need a crawl cycle, and the memory-based engines are on a schedule nobody controls. What you should see instead are leading indicators. Pages getting crawled. A first appearance on a retrieval engine. A source agreeing to a correction.

Month 3. First real movement. Usually where the scorecard starts to move on the questions with the weakest incumbent sources. Not everywhere, and rarely on the head question first. The plan gets re-sequenced around whatever actually moved rather than what was supposed to.

Months 4 to 6. Compounding, or an honest conversation. If the method is working, this is where citation volume bends upward and the earlier work starts reinforcing itself. If it is not, you will have three months of measured evidence to say so with, and the engagement is month to month by then.

Anyone who tells you the first month produces citations is describing a demo, not a program.

Monthly tracking and reporting

End of every month, I re-run the priority questions using the same three-runs-per-engine protocol as the baseline. The scorecard updates. Movement is calculated. The monthly brief has six sections:

  1. Headline number. Where you moved on the priority questions.
  2. Per-piece breakdown. What was published, where it's showing up, what it earned.
  3. Surface insights. Which Reddit thread now cites you, where Perplexity returns you in the top three.
  4. What didn't work and why, written honestly.
  5. Next month's plan, sequenced based on what moved.
  6. Three to five verbatim buyer-language quotes captured during the work.

Briefs are delivered as a written PDF plus a 30-minute Loom walkthrough. The Loom is what builds the relationship. Static reports get skimmed.

What counts as movement, and what is noise

Because engine answers are non-deterministic, a single changed result proves nothing. The rule is simple and it is applied the same way every month: a tier change has to hold on at least two of the three runs to be recorded as movement. One run out of three is logged as noise and left alone. Programs that report from a single run will show you dramatic monthly swings in both directions, none of which are real.

Leading indicators, which is what the early months are judged on

Citation volume is a lagging indicator and it stays flat while the program loads. Judging months one and two on it is how good work gets cancelled early. So the brief separates the two explicitly. Leading: pages crawled and re-crawled, first appearances on the retrieval engines, a source agreeing to a correction, a new third-party page mentioning you at all. Lagging: tier movement on the priority questions, citation frequency, and eventually referral traffic.

The downstream half, and why it undercounts

Analytics only ever shows you a floor. A large share of AI-referred visits arrive with no referrer and land in direct traffic, and the native AI channel Google added to GA4 in May 2026 covers some engines and not others. So the brief reports the analytics number as a floor rather than a measurement, and the citation scorecard carries the actual signal. Anyone presenting AI referral traffic as a complete count is either not aware of this or hoping you are not.

What is deliberately not in the brief

Word counts. Pieces published as a headline figure. Impressions without position. Anything that goes up simply because the month happened. If a number cannot change a decision about next month, it is not in the report.

The adaptation protocol

AI search is not stable. Engines change retrieval logic. Platforms shift moderation rules. Reddit hardens against marketing one quarter and softens the next. The methodology is built to evolve.

First Monday of every month, I run a small set of control prompts across all four engines to detect retrieval-pattern shifts. If the patterns change, the SOP gets updated within 14 days. Surfaces that stop producing citations get demoted. New surfaces that start showing up get promoted. None of this is ad-hoc. The protocol is documented. The protocol changes when the data demands it.

The control set, and why it is not your questions

The monthly control prompts are a fixed set of questions that have nothing to do with any client. They never change, and no work is ever done to influence them. That is the entire point. When a client answer shifts, there are two possible causes: the work moved it, or the engine changed underneath everyone. Running an untouched control set is the only way to tell those apart. Without it, every engine-wide change gets misread as either a win or a failure that belongs to the program.

What the control set is watching for

  • Whether an engine is retrieving live or answering from memory on a given question type. This flips, and it changes what is even reachable in the near term.
  • Whether community sources are being weighted up or down. This has moved in both directions in the last year.
  • Whether the engine names specific vendors at all, or has started describing categories and refusing to shortlist.
  • How many sources it cites, and whether it has narrowed to a smaller, harder set.

What happens when something shifts

The plan gets re-sequenced, not rewritten. A surface that stops producing citations is demoted for the next cycle rather than abandoned, because these things reverse. A surface that starts producing them gets promoted immediately. The change and the reason for it go in that month's brief in plain language, including when the honest answer is that an engine did something nobody has an explanation for yet.

Senior editing plus AI production

AI handles drafting and research. Editing, fact-checking, voice match, and final approval go through me. AI is leverage, not a replacement. This is told to clients openly. Premium buyers respect honesty about AI use more than they respect agencies pretending to be 100% human at $8K - 25K/month.

The senior editing pass is what makes the difference. Generic AI output gets skipped by AI engines just like it gets skipped by readers. Original framing, real expertise, and real attribution are what gets cited.

Where the line actually sits

AI does first-pass research, structural drafting, and the tedious parts of the measurement work. It does not decide what the piece argues, and it does not get the final read.

What goes through me on every piece: the argument itself, because a piece with no position gets skipped by engines and readers equally. Every factual claim, checked to a primary source. The voice, matched to how the client actually talks rather than how a model thinks a B2B company talks. And the decision to cut anything that is technically true but adds nothing, which is most of what a first draft contains.

The rule that costs the most and matters the most

No figure ships unless it can be traced to a primary source. Not a statistic quoted in another article, not a number that appears in three blog posts without any of them saying where it came from, and not a plausible-sounding figure a model produced. This rule removes a lot of confident-sounding sentences from drafts, and it is the single biggest reason the work takes as long as it does.

It exists because unsourced claims corroborate beautifully. Repeat one often enough across the web and the engines start returning it as established fact. A field that publishes its own marketing as evidence is teaching the machines its own marketing, and I would rather not add to it on a client's domain.

What the client approves

Every piece before it publishes, and the monthly plan before the work starts. Nothing ships on your domain that you have not read.

Tool stack

Every tool I use, named publicly. No affiliate disclaimers needed because there are no affiliates.

  • Ahrefs for keyword research and competitor SERP analysis
  • Perplexity Pro for citation source mapping
  • Claude and Claude Code for drafting and research
  • Notion for per-client tracking and documentation
  • Google Search Console and Bing Webmaster for owned-content indexation
  • The four LLM engines (ChatGPT, Claude, Perplexity, Google AI Overviews) for measurement
  • Loom for monthly walkthroughs
  • PayPal for invoicing

What is deliberately not in the stack

There is no AI visibility monitoring platform on that list, and the omission is on purpose.

Those tools are improving and some of them are genuinely useful at scale. But they measure, they do not move, and buying one is the easiest item on a program to complete, which is exactly why it tends to get completed first and alone. A dashboard reporting zero for two months is not a program. It is a subscription that agrees with you about the problem.

At five clients, the measurement runs by hand against a fixed question set, because I would rather spend the hours on the source work than on a tool that tells me what I already found. The point at which a platform starts paying for itself is a capacity signal, not a maturity badge: when manual checking has become the bottleneck. That is not where this operation is, and I will say so rather than add a line item.

The uncomfortable part of a short tool list

Most of this work is manual. Reading the cited pages. Writing to the person who owns a stale comparison article. Checking a claim back to its source. There is no product that does the part that actually moves an answer, and a long tool list would mostly be a way of implying otherwise.

What this methodology will not do

It won't guarantee revenue. Citations and rankings are leading indicators, not directly attributable conversion events. Most clients see pipeline lift from the work. Some don't, because their product or sales motion is the bottleneck. I'm honest about this on every sales call.

It won't replace traditional SEO. Companies winning in 2026 do both. The shorter pieces in the six-piece-a-month rhythm are doing traditional SEO work. The flagships are doing AI search work. They reinforce each other.

It won't work for clients who want 20+ pieces a month. The whole point is depth, not throughput. Anyone whose first question is volume is the wrong fit, and I'll say so on the discovery call.