toolkit

Scenario · The bigger jobs

Which pages are worth keeping?

“We've got two hundred pages and no idea which ones are doing anything.”

Two hundred pages with no idea which ones are actually doing anything is a decision that's impossible to make one page at a time — it has to happen as a single spreadsheet, or it never happens at all. This builds exactly that: a keep, update, or prune recommendation for every page with a stated reason, folding in real search performance so a page with no visible traffic isn't automatically marked for removal — because it might be the one a referring doctor actually reads, and that judgment call always stays with a person.

What to ask for

See it work

A real run of Content inventory:

REVIEW — 1 page(s) · 0 keep · 0 update · 1 prune
  PRUNE      28w  https://toolkit-tests.pages.dev/content-write/inventory-thin

  csv:    ./out/content-inventory.csv
  report: ./out/content-inventory-content-inventory.html
  json:   ./out/content-inventory.json

The captured report, exactly as a run hands it to a client —open the full report ↗

Seen in the wild

75 of this site’s 163 pages have no link into them.

Ticket #12668open

Internal linking across a set of procedure pages

Page content

What was asked for

Please optimize these pages to link to all procedure pages, about us, team, etc.
The request on the ticket, verbatim. It names the pages to work on.

This is a good request and it was being worked correctly — open the pages the client listed, add the links, reply. The only thing nobody had was the shape of the site those pages sit in. So before touching anything, the whole link graph was built: every page the site publishes, and every internal link between them.

What the link graph pulled back

Pages in the graph163
Nothing links to them75
Share of the site46%
Unreachable from home76
Held by one link only33
VerdictFAIL

17 URLs the sitemap declares could not be fetched. They were left out of the graph rather than counted as orphans — an unfetched page is unknown, not proven missing.

What that costs on a site that sells procedures

  • 75 pages Google may never reachNot found
  • 76 a visitor cannot navigate toNot found
  • 33 pages hang off a single linkNot found
  • 27 links that say only click hereUnseen
The pages namedWhat the ticket asked for, done well
163 pagesEvery page the site publishes, mapped once

Content sprawl is invisible until someone counts it. The output here is a spreadsheet with a recommendation per page, which is the only format in which a two-hundred-page decision can actually be made.

  1. Pull the performance data first. Clicks and impressions per page from Search Console over a meaningful window — twelve months, so seasonality does not read as decline. Without this the inventory is a word count and nothing more.
  2. Run the inventory. It gathers word count, freshness and readability per page, folds in the Search Console numbers, and sorts everything into keep / update / prune with a stated reason per row. This is the artefact; everything else here supports reading it.
  3. Read the prune list against reality, not against the numbers. Zero traffic is a signal, not a verdict. Legal pages, location pages, and pages that exist to be linked from an email get no search traffic by design.
  4. Check the update list for the cheap wins. A page with impressions and no clicks is usually a title and description problem, which is hours of work rather than a rewrite. A page that is stale rather than bad is a freshness pass.
  5. [manual] Agree the list with whoever owns the content, in one pass rather than page by page. This is a single conversation about a spreadsheet, and splitting it into two hundred conversations is how the project dies.
  6. [manual] Execute the prunes as removals with redirects, not deletions. Each pruned page follows the page-takedown procedure — a 301 to its parent, references swept. Two hundred bare 404s is a worse outcome than the sprawl.
  7. Re-run the inventory afterwards as the record of what the site is now.

Prune is the row that needs a person

Keep and update are safe to act on from the data. Prune deletes something, and the reasons a page has no traffic are more varied than any tool can see. The inventory is deliberately a recommendation and not an instruction.

What this does not cover

The decision. It produces a keep / update / prune recommendation per page with the reasons attached, and every prune is a judgement someone has to own — a page with no traffic may be the one a referring surgeon reads, or a legal requirement. It also cannot see value that does not show up as traffic or words: a thin page that closes bookings by phone looks identical to a thin page nobody wants.

← All scenarios