Most of your old content cannot be saved. A small part of it is worth everything.

/ 7 min read / By Faz

Every B2B SaaS company that has been publishing for a few years arrives at the same question once AI search starts to matter. There are two hundred posts in the archive. Some of them still get traffic. Almost none of them get cited. So what do we do with all of it?

The advice you will find says refresh. Update the statistics, tighten the headings, add an FAQ block, republish with a new date. There is a number that circulates alongside it, showing that pages cited by AI engines tend to have been updated recently, and it gets used as proof that updating causes citation.

It does not prove that. It is a correlation, and the ordinary explanation fits it better: the pages a company keeps current are the pages it cares about, which are usually its best pages on its most important topics. Recency is a symptom of maintenance, not the mechanism. If you take the finding literally and start pushing dates on two hundred posts, you will have done a great deal of work on the symptom.

I am not saying freshness is irrelevant. On the engines that retrieve live, a stale page with dead facts genuinely loses to a current one. I am saying refresh is the wrong first question, and it is wrong in an expensive way, because it sends you page by page through an archive most of which was never going to be cited under any amount of editing.

The actual problem is selection

Old content fails in AI search for a reason that editing rarely reaches. Most of it was written to be read in sequence. It builds an argument, it sets up context, it pays things off three paragraphs later. That is good writing for a human who chose to read it, and it is close to useless to an engine that needs to lift a self-contained answer to one specific question and attribute it to you.

Which means the archive splits into groups that need completely different decisions, and treating them as one pile is the error underneath the refresh treadmill.

Some pages are structurally fine and merely out of date. The facts have moved, the product has changed, the screenshots are old. These are the cheap wins and they are the ones people mean when they say refresh.

Some pages have a real, defensible answer buried inside them that has never been surfaced. The page ranks for something, someone put genuine thinking into it, but the answer to the question the title asks does not appear until the middle and is never stated cleanly. These are not refreshes. They are rebuilds, and they are where the return is.

Some pages are one of six near-identical takes on the same topic, published over three years by four different people. Refreshing all six is worse than refreshing none of them, and I will come back to why.

And some pages are simply thin. They were written to hit a monthly count, they say what everyone says, and no amount of restructuring gives an engine a reason to prefer them. Editing these is the most comfortable work available and the least useful.

Why refreshing the overlapping cluster makes it worse

This is the part that surprises people, so it is worth being precise.

When an engine assembles an answer, it is choosing sources. If you have six pages circling the same question, none of them is the obvious one. Your internal links are split across them, any external links you have earned are split across them, and none of them accumulates the corroboration that would make it the page on that topic. You have divided your own evidence six ways and then done a round of edits on every piece of the division.

Refreshing all six does not fix that. It produces six recently updated pages that still compete with each other. The move is to pick the one with the strongest claim, fold the genuinely distinct material from the others into it, and redirect them. That is the same consolidation logic that should govern what you publish next, applied backwards to what you already published.

One page that answers the question completely beats six that each answer part of it, and it beats them by a margin no editing pass on the six will close.

A triage you can run in an afternoon

Take the archive and sort every page into one of four buckets. Do not start editing until every page has a bucket.

Merge. Anything that overlaps substantially with another page. Choose the survivor, move the good material in, redirect the rest. Do this first, because it changes the shape of everything else.

Rebuild. Pages where a real answer exists but is not liftable. You are not updating these, you are restating the answer at the top, in a form an engine can quote without needing the paragraphs around it. This bucket should be small and it should get most of your hours.

Leave. Pages that are structurally sound and still true. Change the facts that are wrong and stop. Resist the urge to touch them further because touching them feels productive.

Retire. Thin pages with no defensible claim. Remove them or redirect them to whichever page does have the claim. Nothing is lost here that was earning anything.

The ratio matters more than the process. In most archives I have gone through, the rebuild bucket is small, often under a tenth of the library, and it is where nearly all the available citation sits. Everything else is either cheap or not worth doing.

What editing will not do at all

One honest limit, because otherwise this reads as a promise.

If an engine is answering a question about your category by citing three third-party pages that do not mention you, no edit to your archive changes that answer. The page is not in the pool. This is the ceiling every on-site tactic runs into: your own site is the weakest source the engine has, and improving the weakest source has a low ceiling by construction.

So run the source audit before you commit a quarter to editing. If the answers in your category are being written out of pages you do not own, the archive work is hygiene and the real work is elsewhere. Doing the archive first because it is inside your control is the most common way this budget gets spent on the wrong thing.

What I got wrong

The first time I ran this properly for a client, I did it in the obvious order. Their best-trafficked posts, top down, refreshed one at a time. Structure tightened, answers pulled up, facts updated. It was clean work and I felt good about the pace.

Two months later the citations had barely moved, and when I looked at why, the reason was sitting in the sheet the whole time. Four of the eleven pages I had refreshed were about the same topic. I had spent a third of the engagement making four competing pages individually better, which is not a thing that helps. The engine still had no reason to treat any of them as the answer.

We merged them into one, redirected the other three, and that single page started getting pulled into answers within the next crawl cycle. The material was not new. It had all been there, spread across four pages I had just finished polishing.

Now the merge pass runs first, before anyone edits anything. It is unsatisfying work, it deletes visible output, and it is consistently the highest-return hour in the project.

Where to start

Pull every URL you have published into one sheet. Group them by the question each one is actually trying to answer, not by category or tag. Wherever a group has more than one page in it, you have found a merge before you have read a word of the content.

Then run your real buyer questions through the engines a few times, because a single run tells you nothing, and note which of your pages get cited at all. That gives you the two numbers that decide the whole project: how much of your archive is competing with itself, and how much of it is in the answer pool today.

Almost everyone is surprised by the first number. Almost nobody needs to refresh two hundred pages.

The full sequence is on the methodology page, and if you would rather have the triage run for you before anyone starts editing, that is what the retainer does.

Want this run for your B2B SaaS?

Founding pricing for the first 5 clients. Methodology fully public. Month-to-month, cancel anytime.

Apply to work together

Leave a comment

Your email address will not be published. Required fields are marked *