Fixing Orphan Pages No One Can Find
A page with no internal links pointing to it barely exists, no matter how good it is.
There is a kind of page that exists and does not exist at the same time. It is live on the server. It renders fine. If you paste its URL into a browser, everything works. But no other page on the site links to it, so unless you already know the address, there is no path that leads there. These are called orphan pages, and every site older than a year or two has them, usually more than anyone suspects.
They accumulate the way clutter accumulates, which is to say invisibly and for reasonable reasons. A landing page built for a campaign that ended, still up because nobody remembered to remove it or link it. A blog post published before the blog had an index. A product page for something discontinued, unlinked from navigation but never deleted. A help article written for one customer and mailed as a URL. A redesign that rebuilt the navigation and quietly dropped links to a dozen pages that the old navigation carried. No single decision created the orphans. The structure changed around them, and they stayed where they were.
Why does this matter? Think about how anything gets found on the web. A search engine discovers pages mostly by following links from pages it already knows. A page nothing links to may never be discovered at all, and if it is discovered, through a sitemap file for instance, the engine faces an awkward signal: the site’s own author did not consider this page worth a single link. Internal links are how a site expresses what it thinks is important. Search engines read that structure and distribute weight along it. An orphan receives none of it. It sits at the bottom of the site’s own implied ranking, below every page that got even one link from a footer.
Users are in the same position, only worse, because users do not read sitemap files. A human can only reach an orphan through an old bookmark, an external link, or a search result, and the search result is unlikely for the reasons above. So the page’s actual audience rounds to zero. I find this genuinely wasteful in a way that bothers me. Someone wrote that page. It cost hours or days. The marginal cost of connecting it to the site was one link, and the link never happened, so the whole investment sits inert.
Finding orphans is the interesting part, because by definition you cannot find them by browsing. Browsing follows links, and orphans are the pages links do not reach. You need two lists. The first is every page that exists, which you can get from your CMS, your server, or your sitemap if it is honest. The second is every page reachable by crawling, following links from the homepage outward the way a search engine would. Subtract the second list from the first. What remains is the set of pages that exist but cannot be reached. The first time you run this on a mature site, the length of the list is usually a surprise. This mechanical comparison is exactly the sort of thing that should be automated, and it is one of the reasons I built GazeSite around a real crawl of the site rather than a list of URLs someone claims exist.
Once you have the list, each orphan poses one question: should this page exist? Often the honest answer is no. The campaign ended, the product died, the content went stale. For those, delete the page and return a proper status. If the content moved or has a clear successor, redirect the old URL to the new one with a permanent redirect, so old bookmarks and external links keep working. Deleting is not failure. A site is better for having fewer, better-connected pages than for hoarding everything it ever published.
For the pages worth keeping, the fix is to give them the link they never had, and the craft is choosing where from. A link from a related article, in context, is worth more than a link from a buried index, because it tells both readers and crawlers what the page relates to. The best question to ask is: a reader on which existing page would want this one next? Put the link there. If you cannot answer that question for a page, that is evidence for the delete pile after all, because a page with no natural neighbors probably has no natural audience either.
There is a preventive habit that beats all cleanup, which is making linking part of publishing. Every new page should go live with at least one inbound link from an existing page, decided at publish time, the same way the title is decided at publish time. This is a small rule, easy to keep, and it makes orphans structurally impossible going forward instead of a mess to excavate later.
The deeper point is that a website is not a folder of documents. It is a graph, and a page’s position in the graph is part of the page. Content that is not connected is not really published. It is just stored. The difference between the two is one link, which makes orphan pages one of the cheapest problems in all of SEO to fix, and one of the strangest to leave unfixed.
More articles
Most visitors are comparing you against open tabs, and your page either survives that comparison in ninety seconds or loses by forfeit.
Read →Why so many websites forget to ask visitors to do anything, and what that omission costs.
Read →A page that ends without offering a next step quietly hands the visitor back to their other tabs.
Read →