Where Did That URL Come From? Tracing Website-Generated URLs

Unexpected URLs are clues, not automatically crises. Learn how to trace each URL back to its generator, assess the commercial risk and choose a proportionate fix.

Your ecommerce site has a category page at /running-shoes/. Then a crawl finds hundreds of URLs such as /running-shoes/?filter_colour=blue, /running-shoes/?sort=price-low and /running-shoes?size=10.

Some may be useful filter states. Some may duplicate the main category page. Others may have been created by an old template, a campaign link or a crawler pressing buttons that no customer ever sees.

The useful first question is not “how many URLs are there?” It is “who or what created them?”

That distinction matters because an unfamiliar URL is not automatically a technical crisis. It may be useful, controlled, harmless or commercially risky. Before recommending removal, blocking or a development ticket, you need to understand its source and what it is doing.

This article sets out a source-attribution method: a way to build an evidence trail from an unexpected URL back to the website feature, template, feed, application rule, external system or historical link that produced it.

URL discovery is not the same as URL generation

These terms sound similar, but they describe different events:

  • URL generation is when a system constructs a URL. A filter interface might append ?filter_colour=blue to a category path.
  • URL discovery is when a person, search engine, crawler, feed reader or another system learns that the URL exists.
  • URL requesting is when a browser, bot or other client asks the server or CDN for the URL.

A current website may generate a URL that nobody has discovered yet. An external website may link to a URL that the current site no longer generates. Google may discover a URL through a sitemap, an external link or a page it has already seen, even if there is no current internal link to it.

Crawlable HTML links are an important discovery route, but they are not a complete inventory. URLs can also come from sitemaps, feeds, forms, direct requests and application behaviour. A crawl may therefore show what a particular crawler could reach, rather than every URL the site has generated or that others have requested.

That is why a crawl, sitemap export, Search Console report or server-log extract answers a different question. Each is evidence. None is the whole story. Google’s crawling troubleshooting documentation is useful context for interpreting these reports, but it should not be read as a guarantee that any individual report contains every historical or current URL.

What URL provenance means

URL provenance is the chain of evidence linking a URL to the thing that generated it or the route through which it was discovered.

For example:

  1. A crawler finds /running-shoes/?filter_colour=blue.
  2. The URL appears in a link from a category-page filter component.
  3. The component adds the parameter when a customer selects a colour.
  4. The application returns a filtered product list using the same template as the main category page.
  5. The sitemap contains only the clean category URL, while server logs show Google requesting several colour and sort combinations.

That is a provenance trail. It tells you where the URL came from and what happened afterwards.

A useful provenance record should capture:

  • the URL pattern or URL family;
  • the suspected generator;
  • the discovery route;
  • the request evidence;
  • the page response and content behaviour;
  • the classification;
  • your confidence in the attribution;
  • the commercial impact;
  • the proposed validation plan.

This is different from merging every URL found by a crawl, sitemap and Search Console into one large spreadsheet. That may be useful groundwork, but it does not explain why the URLs exist.

Start with the URL pattern, but treat it as a clue

The URL itself often suggests a possible generator. Query parameters such as sort, filter_colour, q, utm_source or page point towards different kinds of feature. A path such as /search/blue-running-shoes/ may suggest a landing-page route, a CMS rule or an old campaign structure.

For investigation, inspect the domain, path, query and fragment separately. A URI has distinct components, and the query may describe a requested state such as a filter, search term or tracking value. These components provide useful clues, but they do not identify which system originally constructed the URL.

A parameter name is not proof of origin. filter_colour might come from the current filter interface, an old version of the site, a partner feed, a manually constructed link or a bot inventing plausible URLs. The pattern gives you a hypothesis to test, not a verdict.

Group URLs into families before investigating individual examples. For instance:

  • /running-shoes/?filter_colour=blue
  • /running-shoes/?filter_colour=red
  • /running-shoes/?filter_colour=green

These are probably instances of one generator. A thousand URLs may reduce to three or four patterns, each with a different source and risk profile.

A worked example: the retailer with too many shoe URLs

Imagine a retailer whose SEO team finds 18,000 previously unseen URLs. Most use the category path /running-shoes/, followed by combinations of filter, sort and tracking parameters. This is a synthetic example to illustrate the method, not reported client data.

The first temptation is to call the whole set “faceted-navigation URLs” and send a recommendation to block them. That skips the investigation. The URLs may not all have the same source or business value.

1. Inspect a representative sample

Choose examples from each obvious pattern rather than inspecting hundreds of near-identical URLs. Record:

  • the complete URL, including its query string;
  • the HTTP status and any redirect;
  • the page title and main content;
  • the canonical signal, where present;
  • links back to the clean category page or to other filtered states;
  • whether the selected filter visibly changes the products;
  • whether the page can be reproduced through the live interface.

Suppose filter_colour=blue changes the product list and is available through the customer-facing filter. sort=price-low changes ordering but not the products. A URL containing utm_source=affiliate produces exactly the same page as the clean category URL.

These are three different behaviours. Treating them as one problem would make the recommendation less precise.

Google’s URL-structure guidance and duplicate-URL guidance provide useful context here: multiple URLs can lead to identical or near-identical content, and sites can provide signals that help search engines select a representative URL. Those signals do not guarantee that every variant will never be requested or processed.

2. Check internal referring links

Search the crawl data and rendered HTML for links to the URL family. Does the filter component create ordinary links? Does a JavaScript interaction update the address bar? Does a footer, feed module or recommendation widget append a parameter?

In the example, the crawl finds that the colour filters link directly to parameterised URLs. The sort control also changes the URL, but it is implemented as a form that the crawler only reaches when it interacts with the page.

This explains why an HTML crawl found some patterns but not others. It also shows why a crawl cannot be treated as a complete inventory. Form-generated states, client-side interactions and URLs that are no longer linked can remain outside the crawl.

3. Compare the page response and template

Fetch the clean category URL and several variants. Compare the status, redirect chain, page content, canonical element, robots directives and internal links.

You are looking for evidence of a shared template or route. If the filtered pages use the same category template and the product list changes according to the parameter, the filter application is a strong generator candidate.

If a supposedly current URL returns an old page title, a generic error page or a redirect to an unrelated location, the source may be historical residue rather than a live feature. The response can help classify the URL, but it usually cannot prove who originally created it. A server can continue serving an old path long after its original template has disappeared.

4. Check sitemaps, feeds and deployment rules

Look for the URL pattern in XML sitemaps, product feeds, affiliate feeds and other files that publish links. A sitemap normally communicates URLs that a site considers important or preferred. Its presence is not proof that a URL is indexed, and its absence is not proof that the URL is not generated or discoverable. Google’s canonicalisation documentation explains the role of sitemaps and other canonicalisation signals in this context.

In the retailer’s case, the clean category URLs appear in the XML sitemap, but the parameterised URLs do not. That reduces the likelihood that the sitemap generated them. It does not rule out the filter interface, internal links or external discovery.

Also ask what changed recently. A deployment may have altered how filters build links. A feed may have started appending campaign parameters. A CDN or application route may still support a path that the CMS no longer exposes. Release notes and template history are particularly useful when current reproduction fails.

5. Examine logs where available

Server or CDN logs can show whether the URLs were requested, when they were requested, by which user agents, how often and with what response. They can also contain a referring URL through the HTTP Referer field. Google’s crawl troubleshooting documentation and Crawl Stats documentation provide useful context for interpreting request data.

That evidence helps answer “what happened to this URL?” It does not necessarily answer “where was it first created?” Logs measure requests, not the original construction event. They may also be affected by caching, retention limits, separated CDN and origin data, retries, monitoring tools and unreliable crawler identification.

Likewise, a missing referrer should be recorded as unknown. The HTTP specification allows referrer information to be omitted, reduced or altered, so its absence does not prove that somebody typed the URL directly into a browser. Nor does it disprove internal generation: RFC 9110’s Referer field definition.

Suppose the logs show requests attributed to Googlebot for thousands of blue, red and green filter combinations, while the site’s own crawler finds links to only a small sample. That suggests a meaningful request pattern, but it still needs to be joined to the filter component and the page response before you label the component as the generator.

6. Reproduce the URL from the suspected source

The strongest evidence is reproducibility. Select the same filter, submit the same search, follow the same template link or run the same feed process. If the same input produces the same URL and corresponding response, confidence in the attribution rises sharply.

For the retailer, selecting “blue” in the live filter produces /running-shoes/?filter_colour=blue. Selecting “red” produces the matching red URL. The application code confirms that the filter component builds the query string.

That is stronger than noticing that the word filter_colour looks familiar. It links the pattern to a live mechanism.

If you cannot reproduce it, do not force a conclusion. Record the possibility of historical generation, an external link, a removed template, a feed, a campaign platform or a crawler constructing the URL itself. Then use release history, old logs, backlink data and platform owners to narrow the possibilities.

Classify the URL before deciding what to do

Attribution tells you where a URL comes from. Classification tells you what kind of thing it is.

A practical classification might include:

  • Useful page: a distinct landing page that serves a clear customer or commercial purpose.
  • Valid state or variant: a meaningful filtered, sorted, regional or session state that may need to remain available, even if it should not become a primary search landing page.
  • Duplicate path: another URL for substantially the same content, with no strong reason to exist separately.
  • Dead residue: a historical URL that no longer represents a live feature or useful journey.
  • Accidental template output: a URL produced because a component, feed or CMS rule added something unintentionally.
  • Crawl trap: a generator that can create a large or effectively endless set of URLs with little additional value.

The classification should follow the evidence, not the URL’s appearance. A parameterised URL is not automatically wasteful. A clean-looking path is not automatically useful.

In the worked example, colour filters may be valid states because they help customers find products. Sort parameters may be low-value variants because they reorder the same products. Affiliate tracking parameters may be useful for attribution but poor candidates for search discovery. The action for each pattern should therefore be different.

Use a scale-and-impact test

Once you know the likely generator and classification, assess whether the issue is worth fixing now.

Consider:

  • Generation scale: how many variants can the mechanism create, and how quickly can that number grow?
  • Crawl activity: are bots requesting the URLs, or did one tool simply find a handful?
  • Indexation: are unwanted variants appearing in search results or being reported as indexed?
  • Duplication: do they repeat content and compete with more useful URLs?
  • Server and platform demand: do requests consume meaningful application, database, CDN or crawl resources?
  • Internal prominence: does the site repeatedly link to the variants from important pages?
  • Commercial displacement: could these URLs cause valuable category or product pages to be crawled, indexed or selected less effectively?
  • Fix cost and risk: how much engineering effort is required, and could the change damage filtering, tracking or product discovery?

There is no universal number at which an unfamiliar URL becomes an emergency. A small number of prominent, indexable duplicates may deserve attention sooner than thousands of harmless campaign URLs. Conversely, a large count may be less important if the variants are not requested, not indexed and not consuming meaningful resources.

This is an interpretation and prioritisation framework, rather than a published search-engine threshold. Google’s Crawl Stats documentation and crawling troubleshooting guidance can help you understand request activity and crawl problems, but they do not provide a single commercial-risk score for every website.

Record confidence, not just conclusions

A provenance investigation becomes much easier to review when it records how certain each conclusion is.

For each URL family, write something like:

  • Pattern: /running-shoes/?filter_colour=*
  • Suspected generator: category-page filter component
  • Discovery route: internal HTML links and crawler interaction
  • Request evidence: repeated requests attributed to Googlebot in CDN logs
  • Response: 200 status, filtered product list, shared category template
  • Classification: valid state or variant
  • Confidence: high, because the live interface reproduces the pattern
  • Commercial impact: medium; useful to customers but capable of generating many combinations
  • Validation plan: test a source-level change on representative filters, then monitor navigation, requests and important category visibility

Also record what would disprove the hypothesis. If removing a filter selection stops producing the URL, that supports the attribution. If the same pattern continues through an affiliate feed or external referrer, the filter may not be the only source.

Remediate the source, then validate the result

Where practical, fix the generating feature or rule rather than repeatedly cleaning up its output. That may mean changing a template, feed, application route, CMS setting or deployment rule. The correct technical control depends on the classification and the intended user journey.

Google treats redirects, canonical annotations and sitemap inclusion as signals when selecting a representative URL, rather than as universally absolute commands. They can form part of a controlled solution, but they do not replace understanding the generator: Google’s duplicate-URL guidance and canonicalisation guidance.

After implementation, validate two things:

  1. the unwanted pattern stops being generated or materially reduces;
  2. important customer and search journeys still work.

Re-run representative crawls, check logs and sitemaps, inspect key templates and monitor search reporting over an appropriate period. A change that removes 20,000 URLs but also removes useful filter navigation is not a successful fix. SEO has a talent for turning tidy spreadsheets into awkward product decisions.

When specialist help is worthwhile

An unfamiliar URL does not, by itself, require an SEO project. An in-house team can often trace a small, well-understood pattern by inspecting the interface, templates and a few log entries.

Specialist support becomes more valuable when evidence conflicts, the site has several CMS and application teams, the generator may sit across CDN and backend rules, historical evidence is important, or the scale makes manual validation unsafe. The difficult part is usually not finding another URL. It is deciding which source is responsible, which variants matter commercially and how to change the source without breaking useful journeys.

Liquid Silver can help map those relationships across crawls, templates, logs, sitemaps, feeds and search data, then turn the findings into an implementation plan. Our technical SEO services cover diagnosis and prioritisation, while our SEO implementation support helps teams move from evidence to a tested change.

The useful question is not “is this URL strange?”

Unexpected URLs are evidence of something. They may reveal a useful customer feature, a historical system, an external discovery route or a generator that is quietly multiplying low-value pages.

The practical method is to separate generation, discovery and requests; follow the URL pattern through its referring links, response, sitemap, feed, logs and current templates; reproduce the suspected mechanism where possible; then classify the result and weigh its commercial impact against the cost and risk of fixing it.

That provenance trail gives you a better answer than a raw URL count. It tells you what created the URL, what the URL is doing now and whether the sensible response is to retain it, control it, monitor it or fix the source.

Share this article

Found this useful? Pass it on.

Share on LinkedIn · Share on X