Migration residue: how to trace legacy URLs after a site move

Legacy URLs can keep appearing after a migration even when redirects are working. Learn how to identify the live or historical source exposing each URL, prioritise remediation and validate the result over time.

A site migration can be technically successful while old URLs continue to appear in crawl reports, Search Console, backlink tools and server logs. That does not, by itself, mean the redirect plan has failed.

The more useful question is: what is still exposing this particular legacy URL?

A current internal link, sitemap entry, canonical annotation, hreflang reference or feed calls for a different response from an old external backlink or a search engine revisiting a URL it learned months ago. This guide treats each legacy URL as an evidence object. It explains how to identify current and historical sources, prioritise what can be changed and validate source removal separately from the slower decline of historical requests.

Separate discovery, crawling, indexing and serving

Search systems do not treat “the URL appeared in a report” as a single event. Google describes separate stages for discovering URLs, crawling them, indexing their contents and serving them in search. A URL can be known but not yet crawled, crawled but not indexed, or indexed without being served for a particular query. See Google’s explanation of how Search works.

After a migration, the distinction helps separate different types of evidence:

  • Discovered or known: a search engine or tool has learned that the URL exists. This does not prove a recent live link.
  • Crawled: a bot requested the URL. Logs can often confirm this, although they may not reveal how the URL was originally found.
  • Indexed: the URL, or information associated with it, has been considered for a search index. A redirecting URL is not necessarily indexed as a standalone result.
  • Served: the URL or its associated destination appears in search for a user query.
  • Reported: a monitoring platform, crawler or backlink database has recorded the URL. The record may be current or historical.

A legacy URL in a crawl export or Search Console report is therefore a lead for investigation, not proof that the current site still links to it. The timestamp and type of evidence matter.

Make the URL an evidence object

Do not begin by looking for one universal migration error. Select a representative legacy URL or URL pattern and create an evidence record for it.

Record at least:

  • the legacy URL and its current status code;
  • the full redirect chain, including protocol, host, path and query-string changes;
  • the final destination and whether it is relevant;
  • where the URL was observed and when;
  • whether the source appears current, historical or uncertain;
  • the likely owner of the remediation; and
  • the next validation check.

A useful record might say: “https://example.com/old-guide returns a direct 301 to /guides/new-guide. It remains in the HTML of three current pages and in the RSS feed. Search Console also reports it, but that report does not establish a current source.”

This avoids treating every appearance of a legacy URL as equally urgent. A current, repeated first-party reference is usually more controllable than a historical third-party observation. That is an operational judgement, not a Google ranking rule, and it should be adjusted for referral traffic, search significance, infrastructure cost and user impact.

Trace the possible discovery sources

Internal links in HTML

Internal links are a sensible first check because the site controls them and templates can repeat them across navigation, related-content modules or page content. Google explains that crawlable links help it find pages, although the existence of a link does not guarantee a particular crawl event.

Run a crawl that exports both the source page and the linked legacy URL. Then inspect more than the rendered report:

  • crawl the raw response HTML;
  • inspect the rendered DOM after JavaScript execution;
  • check navigation, breadcrumbs, pagination, related-content blocks and footer components;
  • search CMS content, structured content fields and template repositories for the old path;
  • check links generated through APIs or client-side components; and
  • include non-production environments if they are publicly accessible or being crawled.

Comparing raw HTML with the rendered DOM matters where JavaScript changes the links available after page execution. If the legacy URL is absent from the server response but appears after rendering, a raw-only crawl will miss an active first-party source. A string in a script bundle, on the other hand, is not automatically an ordinary crawlable link.

Fix the generating source rather than adding another redirect. If a shared template creates thousands of references, it will usually outrank an isolated editorial link. Re-crawl the affected template and a sample of pages that use it.

XML sitemaps

Inspect every sitemap index and child sitemap, including regional, image, video, news and automatically generated files. Google states that it can discover URLs through sitemaps, but describes sitemap submission as a hint rather than a guarantee. Sitemap URLs should generally represent the site’s preferred canonical URLs. See the sitemap documentation.

For each legacy URL, check:

  • whether it is explicitly listed;
  • whether a sitemap generator is rebuilding it from stale CMS data;
  • whether it appears in a regional or specialised sitemap rather than the main index;
  • whether the sitemap is cached or served differently at the edge; and
  • the file’s last-modified information and the time at which it was fetched.

Removing the URL from a sitemap removes one first-party discovery signal. It does not guarantee that requests or reports will stop if internal links, external links, redirects or historical knowledge remain.

Canonical annotations

Search the raw HTML and rendered output for rel="canonical", then inspect HTTP headers where non-HTML resources are involved. A legacy URL in a canonical annotation is clear evidence that the current site is publishing that URL as a signal. It does not prove that Google selected it as the canonical.

Google describes canonical declarations as signals rather than absolute commands and may select a different canonical during indexing. The relevant references are Google’s canonicalisation documentation and its guidance on consolidating duplicate URLs.

Record the page publishing the stale canonical, the delivery method and whether the legacy URL also appears in a sitemap or redirect rule. Correct the canonical-generating logic, then re-test a representative set of page types. A live inspection alone cannot show that Google’s selected canonical has changed. That is an indexing outcome and needs separate monitoring.

Hreflang annotations

For international sites, inspect hreflang in three places: HTML, HTTP headers and XML sitemaps. Google documents all three methods and recommends reciprocal alternate references, including a reference to the page itself, within the relevant language or regional cluster. See Google’s guidance on localised versions.

A stale alternate such as hreflang="en-gb" href="/old-guide" is an active first-party reference. It may contribute to continued discovery, but its presence does not prove that it caused a crawl or that Google indexed the old URL. After remediation, validate the complete cluster: each current page should point to the intended alternatives, and reciprocal references should no longer nominate the legacy path.

Redirects

A redirect is a response to a request, not evidence of how that request originated. An old URL may be requested because of an external backlink, a search engine’s historical knowledge, an old feed entry or a current internal link. The redirect tells the requester where to go next, but usually cannot identify the original discovery source.

For migration URLs, check that:

  • the response is a suitable permanent redirect;
  • the chain reaches the final relevant destination directly;
  • there are no protocol, host or path loops;
  • the final destination is available and is not redirecting unnecessarily; and
  • query parameters and fragments are handled deliberately.

Google recommends permanent redirects for site moves, avoiding redirect chains and retaining redirects for at least a year. Its site-move guidance is particularly relevant here. The one-year recommendation is not a promise that every historical request or report will disappear within that period.

Do not remove a healthy redirect simply because the URL continues to receive requests. A correctly maintained redirect may be the appropriate response to an old external link that cannot be updated. The remediation target is a broken, irrelevant or inefficient redirect, not the existence of every historical request.

Feeds and other machine-readable outputs

Check RSS, Atom, news, product, media and other feeds that may be generated separately from the main page templates. Inspect entry links, identifiers, permalinks and enclosure-related URLs. These are separate machine-readable output surfaces and can expose old URLs even when the visible site appears clean. Google’s description of how information is organised also recognises feeds and other sources as ways information can be made available to systems.

A feed containing a legacy URL is evidence of an active first-party output. It is not proof that Google discovered the URL through that feed; that claim needs corroboration from access logs or other evidence. Inspect the feed source in the CMS, its cache headers and any syndication or feed-management layer. Fetch the feed again after deployment and confirm that old entries have aged out or now contain the intended current URLs. For background, see Google’s overview of how information is organised.

External links

Use Search Console’s Links data, backlink platforms and referral analytics to identify external pages that may still point to the old URL. Treat these datasets as leads, not complete and current link graphs. Google notes that its Links report is sampled and may include historical links, grouped URLs or links that have since been removed; see its documentation on crawlable links and the Links report.

Where an external link is important, current and realistically editable, request an update. Prioritise publishers sending meaningful referral traffic, high-value partners and pages that are frequently crawled. Do not spend equal effort trying to remove every low-value or inaccessible historical reference. A direct, relevant permanent redirect remains necessary because some external publishers will never update their links.

Access logs can show that a legacy URL was requested and how the server responded. They may include a referrer, but that value can be absent, stripped or misleading. Logs therefore corroborate activity; they rarely prove the original discovery route. Google’s crawling troubleshooting guidance provides useful context for interpreting request data, although logs may also contain users, monitoring systems and non-search bots.

Cached and historical references

Old crawl exports, migration snapshots, Search Console observations, monitoring alerts and backlink databases can preserve a URL after the live source has gone. Google’s current documentation says that its former cached-result link and cache search operator are no longer available as current diagnostic tools. For historical investigation, use dated exports, old crawls, CMS records, logs and third-party datasets instead; see Google’s Search updates.

Label this evidence accurately. “Observed in a crawler export from March” is different from “currently linked from the site”. Historical evidence can explain why a URL remains known, but it should not trigger unnecessary changes to an otherwise clean migration.

A provenance workflow for each URL or pattern

Work through this sequence for each representative URL or pattern:

  1. Confirm the live response. Record status, headers, redirect hops, final destination and response timing from more than one access point where relevant.
  2. Search first-party inventories. Query crawl exports, raw and rendered HTML, templates, CMS records, sitemaps, canonicals, hreflang, feeds and other generated files.
  3. Correlate with requests. Use server or edge logs to establish whether the URL is being requested, by whom, when and at what volume. Separate search crawlers, users, monitoring tools and other bots where the data allows.
  4. Review external evidence. Compare Search Console, backlink platforms, referral analytics and live referring pages. Note differences in timestamps and URL normalisation.
  5. Classify the source. Mark it as active first-party, active third-party, historical or unresolved. A URL can have more than one classification.
  6. Assign remediation and priority. Consider controllability, repeatability, scale, crawler accessibility, user impact, search significance, infrastructure cost and effort.
  7. Capture a baseline. Save source counts, representative URLs, redirect behaviour and request volume before changing anything.

This process cannot always identify one unique source. It does separate evidence of a current signal from evidence that the URL is simply still known.

Worked synthetic example: one legacy URL, five sources

Consider a fictional travel publisher, Lumen Trails, which moved its international guides from /old/italy/rome-weekend to /guides/rome-weekend. The old URL returns a direct permanent redirect to the new guide, yet it continues to appear in monitoring data.

The investigation finds five different signals:

  • A global “popular guides” component still links to the old path on 420 pages. This is an active, repeated first-party source.
  • The en-GB sitemap contains the old URL because its generator reads a stale CMS record.
  • The French version’s hreflang annotation points to the old English URL. This is an active first-party signal, although it does not prove canonical selection or a crawl caused by hreflang.
  • An RSS feed contains the old URL in an entry published before the migration. Several aggregators still reference that entry.
  • A travel association links to the old URL. The publisher cannot edit the page quickly, but the redirect is direct and relevant.

Lumen Trails should not begin by removing the redirect. It should first fix the shared component, correct the sitemap generator and repair the hreflang cluster. It should then regenerate the RSS feed and ask the association to update its link when practical. The external reference remains acceptable in the short term because the redirect preserves the route to the relevant guide.

After deployment, the team re-crawls the affected templates, fetches all sitemap and feed files, checks the hreflang cluster, reviews redirect chains and compares logs with the baseline. If requests continue but the first-party source inventory is clean, the remaining activity is more plausibly attributable to the association, aggregators, users or historical search-engine knowledge. That is a different problem from an unresolved site-wide template defect.

Validate remediation in layers

Validation should test the source that was changed, not just whether the URL still appears somewhere.

Immediate checks

  • Re-crawl the affected page templates and representative URL sets in both raw and rendered modes.
  • Search current HTML, headers, XML files, feeds and CMS output for the legacy URL.
  • Fetch sitemap indexes and child files directly, checking caches and generated variants.
  • Validate the full hreflang cluster and canonical output.
  • Test the legacy URL and every redirect hop, confirming a direct route to the intended destination.
  • Check deployment and edge-cache timestamps so an old response is not mistaken for current output.

Short-term and time-based checks

  • Compare the number and distribution of first-party references with the pre-remediation baseline.
  • Compare legacy requests in server or edge logs by date, user agent and response status.
  • Review representative Search Console URL states and crawl observations, recording that these may lag behind live changes.
  • Recheck important external referring pages and referral traffic.
  • Repeat the comparison at intervals suited to the site’s crawl rate and migration scale rather than assuming that a fixed deadline proves completion.

The desired outcome is not necessarily zero requests. It is the removal of controllable, repeated first-party exposure; stable, relevant redirects for unavoidable historical references; and a declining or explainable pattern in residual requests and reports.

What migration residue tells you

Legacy URLs are diagnostic clues, not automatic failure notices. Persistent first-party links, stale sitemaps, incorrect canonicals, broken hreflang clusters and feeds deserve direct remediation because the site is still publishing those signals. A direct redirect receiving occasional requests from an old external link may simply be doing its intended job.

The practical distinction is provenance. Establish whether the current site is generating the URL, whether an external source is still sending people or crawlers to it, or whether a tool is reporting historical knowledge. Then choose the response that fits the evidence.

Share this article

Found this useful? Pass it on.

Share on LinkedIn · Share on X