← all articles

What to do when you lose 334 articles

seo technical-seo site-audit content-strategy

334 published articles went missing off a site I own. Indexed, linked to, listed in the sitemap, and then 404.

I found it during an audit I was running for an unrelated reason, re-crawled to make sure I had not fat-fingered a filter, and got the same number back.

The text survived. Every article on that site begins as a source file in a repository, so the words were sitting in version control the whole time. The published copies were what vanished. That gap between “the content exists” and “the site serves it” is the whole story, and it is where the recovery work lives.

The decline that looked like seasonality

The site had analytics. I look at analytics. That is the uncomfortable bit.

Pages do not all die on the same day, and even when they do, traffic drains over weeks while the crawler works through them. What you see is a line sloping gently down. Ten percent one month. Another twelve the next.

A line like that has four ready explanations before it has the right one. Seasonality. The update everyone was posting about. A competitor who finally started shipping. Your own thin publishing month. I have used every one of those to explain away a decline, and so has anyone who has run a site for more than two years.

None of the available explanations was “a third of the pages stopped existing”, because that is not a thing analytics is built to say.

Count your pages, not your sessions

The number that would have flagged this on day one is the count of live published pages.

Not sessions. A plain integer: how many URLs does this site currently serve. Fetch the sitemap, count the entries, write it down.

Nothing in a standard stack tracks that. Search Console reports indexed pages, which is close but lags by weeks and blends real losses into crawl noise. Analytics counts what got visited, which is a subset that shrinks for a dozen innocent reasons. Your CMS counts what it still has, and what it still has is the damaged version.

So the check has to be yours. A sitemap fetch and a line count, on a schedule, with an alert if it moves more than a few percent. That is the entire engineering cost of catching this the day it happens, and I had not spent it.

There is a cheaper detector too, and I had that one running and was not reading it. Your server’s 404 log shows requests for dead URLs, with referrers attached. Scanner noise hits URLs you never had. A run of 404s on URLs you definitely did have, arriving with real referrers, is your own site reporting the loss to you in plain text.

Reconstructing a list of what you had

When you do find out, the reflex is to start restoring. Give it an afternoon first.

You cannot restore a list you do not have, and the site itself will not give you one. Your sitemap describes the damaged version. So does your CMS. Neither knows about pages that are already gone.

You rebuild the list from outside:

  • the search engine’s index, which still holds URLs it has not got round to dropping
  • your analytics page report, going back as many years as retention allows, since every URL that ever got one visit is in it
  • any old crawl export sitting on disk, which names pages that no longer answer
  • public archive services, which have hit your site more often than you would expect
  • your own backlink data, which lists destination URLs on your domain whether or not they still resolve

Deduplicate that into one column and you have a candidate list. It will be wrong at the edges. It is still better than anything the live site can tell you.

Which is the lesson arriving early: keep a dated export of your own URL list. One file, monthly, in the same repository as the content. It costs nothing and it turns a week of archaeology into an afternoon of diffing.

Where the content still lives

Then go looking for the words, in this order.

Version control first. This is why I still have those articles. Source files were committed, so text, metadata and internal links all sit in history at a known commit, and recovery is a checkout rather than a rescue.

Backups second, with the caveat that trips people. A backup only helps if the loss is newer than your retention window. If the pages went four months ago and you keep 30 days, every backup you hold now faithfully contains the damaged site. Plenty of operators learn how their backup policy actually works on precisely this day.

Search engine cache third, and move fast, because it expires. It gives you a rendered copy as of the last crawl.

Archive services fourth. Coverage is patchy and skewed toward pages that had traffic or links, which is normally frustrating and here is convenient, because those are the pages worth having back.

If none of that lands, rewriting is on the table. Sometimes it is the right answer.

Here is the part I would tell anyone to do differently from their instinct.

Do not restore everything, and do not restore chronologically.

On a site with a few years of history, a large share of the archive earns nothing. No traffic, no links, no rankings, published because somebody had a quota that quarter. Putting those back is effort you spend to feel restored rather than to be restored.

Triage on external evidence instead. Three questions, and one yes is enough:

  • does anything link to this URL
  • did it receive search traffic in the last twelve months
  • does it rank for a term you actually want

Then sort the yes pile by inbound links and start at the top, because a 404 on a page somebody linked to is the most expensive dead page you can own.

I sell placements, so I have the numbers on both sides of this.

A dofollow placement on a decent site runs somewhere in the low hundreds. I have bought links, and I have watched at least one I paid for end up pointing at a URL that no longer answered. That is money spent on a link into a hole, and the seller did nothing wrong.

Earned links are worse to lose. Somebody read the piece, decided it was worth citing, and put their own name beside yours. That is harder to acquire than the paid kind and impossible to reorder.

Either way, the link is still live, still pointing, and every click lands on an error. Whatever authority it carried goes nowhere.

The part people underrate is the social cost. The person who linked to you will notice eventually, or their broken-link checker will, and the link comes out. Once it is out, fixing the page three months later does not bring it back. You would have to ask, and asking spends goodwill you would rather keep for something else.

Redirects that are honest, and the ones that hide the problem

Restoring properly takes time, so redirects become the fast patch. This is where it gets slippery.

A redirect is honest when the destination answers the same question. You lost a guide, you have another guide to the same thing, you send the traffic there and say so on the page. The reader gets what they clicked for.

A redirect is dishonest when every dead URL points at the home page. I understand the reflex, because it drives the 404 count to zero and the 404 count is what is being measured. But somebody clicked a link about one specific topic and landed on a front page with no explanation, which is worse than an error page. The error page at least told the truth. Search engines have treated bulk home page redirects as soft 404s for years, so it does not even buy the ranking it was reaching for.

My rule: redirect when I can name the destination out loud and defend it. Otherwise leave the 404 standing until the real page is back. A visible problem gets fixed and a hidden one does not.

Three habits, all of them nearly free

Keep content in version control. If you are on a CMS, export on a schedule and commit the export. The value is the diff, not the backup. A commit deleting 300 files is loud in a way a quiet database change never is.

Count live pages weekly and alert on the delta. Sitemap fetch, line count, one row in a log.

Check the pages other people linked to. That list is short and it is the highest-value list you own. Monthly is plenty.

I had every tool and still missed it

I had the content in version control, which is the only reason any of this was recoverable. I had years of page-level analytics. I had crawl exports on disk. I buy and sell link placements, so I look at link data more than most people ever will.

And I still went months without seeing it, because every one of those tools was pointed at the same question: how is traffic doing. Not one of them was asked whether the site still had the pages it had last month.

Owning the data and watching the data are separate jobs. Recovery took a few days once I knew. The not knowing ran far longer than the fixing did, and that part is nobody’s fault but mine.

The sitemap-count script and the restore checklist I used are here.

for SEOs
Tracking rankings or scraping SERPs at scale?

Rank checkers and SERP crawlers get blocked and geo-skewed fast on datacenter IPs. Singapore Mobile Proxy runs real 4G/5G mobile IPs that search engines still trust, so your position data stays clean.

see plans →
read on
More from The SEO Desk

Technical SEO, link building, content and SERP strategy, and tool reviews for people who ship growth.

browse all articles →