Sitemap hostname drift: trace the origin to its source
A provenance-led method for finding whether incorrect sitemap origins came from deployment variables, generators, reverse proxies, build artefacts or caches.
An XML sitemap can be validly formatted and still contain the wrong origin. Its <loc> values may use a staging hostname, an old domain, http instead of https, an internal port or a host derived from a reverse proxy. The visible symptom is straightforward: the sitemap lists URLs that do not match the intended production architecture. The difficult part is establishing which layer created or preserved that value.
This is sitemap hostname drift: a sitemap URL uses an unintended host, protocol, port or proxy-derived origin. It is narrower than a general set of stale URLs or a deliberately multi-domain site. The useful question is not only “which sitemap URLs are wrong?” It is also “where did the emitted origin come from, and does that value survive the complete production delivery path?”
This article sets out a source-tracing method. It starts with an approved-origin model, follows the hostname through generation modes, application settings, deployment variables, reverse-proxy headers and caches, then finishes with production checks for indexes, partitions and affected URL cohorts.
Start with an approved origin, not a hostname pattern
Before classifying a hostname as drift, define the origins that are intentionally allowed. A single-site implementation may permit only https://www.example.com. An international or multi-brand architecture may also permit regional, language, tenant or brand hosts. Document those hosts as an approved origin set rather than inferring them from a pattern.
This matters because scheme, host and effective port are components of HTTP origin identity. An http URL, an https URL and a URL on a different port should therefore be counted as distinct variants, even when their paths are identical. See the definition of origin in RFC 9110.
Sitemap locations are expected to be fully qualified URLs rather than relative paths, according to both the Sitemaps protocol and Google’s sitemap documentation. That makes the origin visible in every entry, but it does not explain why the origin is wrong.
Build an evidence model before changing configuration
Capture the sitemap output as evidence before editing environment variables or proxy settings. The initial comparison should include:
- the intended production origin and approved alternate origins;
- every discovered sitemap index, sitemap partition and relevant alternate sitemap path;
- the host, scheme and port in each index reference and URL entry;
- canonical URL values on a sample of affected pages;
- internal links from the same pages;
- the HTTP response, redirect chain and final URL for representative sitemap entries; and
- response headers, cache metadata, deployment identifiers and body hashes where available.
The purpose is to separate a sitemap-only problem from a wider origin problem. If canonical URLs, internal links and final response URLs all use the approved production origin while the sitemap alone contains a staging host, the sitemap generation path becomes the leading area for investigation. If every signal uses an alternate host, the issue may be broader than sitemap generation.
This comparison is an applied diagnostic inference from the different roles of sitemaps, canonical URLs and redirects. Google describes sitemaps as a discovery signal and states that it attempts to crawl URLs as listed. Its duplicate-URL guidance gives stronger canonicalisation weight to redirects and rel="canonical" than to sitemap inclusion. Those documents do not prove that a particular wrong-host configuration will cause ranking loss or deindexation. See Google’s sitemap guidance and Google’s duplicate-URL guidance.
A sitemap index and the files it references are separate evidence surfaces. The protocol allows an index to reference multiple sitemap files, and an individual sitemap file is subject to limits of 50,000 URLs or 50 MB uncompressed. Validate the index references and the contents of the referenced files independently rather than treating a clean root response as proof that the complete graph is clean. These limits and the index structure are documented in the Sitemaps protocol.
Classify the symptom, but do not infer the cause
Count variants by scheme, hostname and port. Useful classifications include:
- the approved production origin;
- a staging or preview hostname;
- an internal service name or private hostname;
- an old domain retained from a migration;
- an alternate but deliberately supported regional, language or brand domain;
httpwhere production is intended to behttps;- a non-standard or internal port; and
- a proxy-derived host that varies by request or edge route.
These categories describe the output, not its provenance. A staging hostname could have come from a production environment variable, a hard-coded sitemap base URL, an old static artefact or a cache. An apparently internal hostname could instead be generated from a forwarded host that the application trusts. A second public domain could be entirely valid.
Pattern matching alone is therefore unsafe. A sitemap body containing staging.example.internal does not prove that a staging variable caused it. The hostname is evidence of an emitted value; the root cause must be established by tracing the generation path and reproducing the value at the relevant layer.
Identify how the sitemap is generated
Generation mode determines where provenance can be found. Establish whether the sitemap is:
- built into a static file during deployment;
- generated by a scheduled job and stored in object storage;
- created at request time by an application route;
- generated at the edge; or
- served through a cache in front of one of these mechanisms.
Framework documentation illustrates why these paths should not be diagnosed identically. For example, Next.js sitemap documentation describes sitemap generation through framework conventions that can produce output as part of an application build or route. The example is not a universal taxonomy, but it demonstrates the key distinction: a value fixed in a build artefact requires a different investigation from a value derived from the live request.
Record the implementation details that affect provenance:
- the source file, route or library responsible for sitemap generation;
- the build command or scheduled job that creates the file;
- the variables available at build time and runtime;
- the configured public, canonical or base URL;
- the storage location and object version, if the sitemap is static;
- the cache layers between the public URL and the generator; and
- the headers or request fields used to construct absolute URLs.
Trace the origin through the configuration layers
1. Deployment environment variables
Compare the value used by the sitemap build or runtime process with the value intended for public production. Look for variables named APP_URL, SITE_URL, PUBLIC_URL, BASE_URL, CANONICAL_HOST or their platform-specific equivalents. Do not assume that a variable with a production-sounding name is the one the generator reads.
Check which value existed when the sitemap was built. A corrected deployment variable cannot change an already generated static file unless the build or generation job runs again. Compare the deployed artefact with the source configuration and, where possible, record a build identifier alongside the sitemap body hash.
2. Application and base-URL settings
Inspect application-level URL generation separately from general site configuration. A CMS may have a public site URL, while a sitemap package has its own base URL or may use a library default. A canonical URL setting may be correct even though the sitemap generator reads another setting.
Test the generator directly where possible. Supply a controlled page path and observe whether the output changes when the configured base URL changes. If the direct generator continues to emit the old host, the proxy is unlikely to be the immediate source. This is a diagnostic test, not proof that no other layer is involved.
3. Sitemap-generator logic
Search the generator code, package configuration and templates for hard-coded schemes, hosts, ports and string concatenation. Check whether different sitemap cohorts use different configuration. Product partitions, editorial files, locale indexes and image or video sitemaps may not share the same origin constructor.
Inspect fallback behaviour as well. A generator may prefer an explicit base URL but fall back to request context, a framework configuration value or an environment variable when the explicit setting is absent. That fallback can explain why only one deployment mode or sitemap route is affected.
4. Reverse-proxy forwarded headers
Request-time generators may construct absolute URLs from the request’s host and scheme. In a typical proxy chain, TLS may terminate before the application receives the request, and the proxy may rewrite the upstream Host. The Forwarded header specification describes how original host and protocol information can be carried across proxy hops.
Misconfigured trusted-proxy handling can therefore produce an http sitemap, an internal upstream host or a host belonging to the edge request rather than the public site. Framework documentation gives stack-specific examples: for instance, Laravel’s request documentation explains how trusted proxies affect the scheme and host information exposed to an application.
Test the full chain rather than changing trust settings immediately. Compare:
- the public request received at the edge;
- the headers forwarded to the application;
- the host and scheme seen by the generator;
- the origin emitted in the sitemap; and
- the result when the request is made directly to the origin, bypassing the edge where that is safe and possible.
Forwarded headers are not automatically authoritative. Their meaning depends on which proxy overwrites or strips them and which values the application is configured to trust. This article is concerned with tracing their effect on sitemap output, not with providing a general reverse-proxy security design.
5. Cache and edge delivery
A cache can preserve an old sitemap after the generator has been corrected. HTTP caching allows previously generated representations to be served according to cache rules, and behaviour can differ by provider, host, path, query string and request headers. The relevant semantics are described in RFC 9111.
Compare a normal public request with a cache-bypassed request and, where possible, a direct origin request. Capture cache status headers, age, last-modified values, deployment identifiers and body hashes. If the origin returns the corrected host but the public response retains the old one, stale delivery becomes the leading explanation. If both responses contain the wrong host, continue tracing the generator and its inputs.
A synthetic tracing example
Consider a retail site whose public origin is https://shop.example.com. Its sitemap index is clean, but one product partition lists https://catalog-web-7.internal/products/blue-jacket. The product page’s canonical URL and internal links use the public origin, and the internal hostname redirects when requested through a permitted network route.
The first conclusion should be limited: the partition emits an unintended origin. It should not be assumed that the application’s main site URL is wrong.
The team then compares three paths:
- The static build artefact already contains the internal hostname. The partition generator is using an environment variable populated by the deployment service.
- The build artefact contains the public hostname, but the live request returns the internal hostname. The route is request-time generated and constructs URLs from the upstream
Host. - The origin response is corrected, but the public edge response contains the internal hostname. The remaining fault is stale cache delivery or an unpurged object.
These paths produce the same visible symptom but require different remediations. The example is synthetic: it demonstrates why output classification must be followed by provenance testing.
Scale validation across the complete sitemap graph
Once the likely emitting layer is known, turn the diagnosis into deterministic checks. A production validator should:
- Fetch the root sitemap and discover all sitemap-index references.
- Resolve nested indexes and partition URLs, including compressed files and documented alternate sitemap paths.
- Parse every
<loc>value and count scheme, hostname and port variants. - Compare each variant with the approved origin set.
- Record which index or partition emitted each incorrect value.
- Sample or crawl affected URL cohorts to verify status codes, redirects and final response URLs.
- Compare representative pages with canonical and internal-link origins.
- Run the checks against the fresh build or generated artefact before deployment.
- Repeat them through the public live request path after deployment and cache propagation.
For very large systems, complete parsing is preferable when operationally feasible. If documented sampling is necessary, sample by partition, locale, template, generation job and hostname variant rather than taking only a random sample from the root index. A clean root index does not establish that every nested partition or alternate public sitemap path is clean.
Keep the output suitable for release monitoring. At minimum, retain the approved-origin result, counts by variant, affected file names, representative URLs, response and redirect results, build identifier, response hash and cache status. A failure should identify the partition and emitting path, not merely report that “the sitemap is invalid”.
For the wider mechanics of testing sitemap changes before and after release, see our XML sitemap release validation guide. For large sites where the investigation involves partition coverage and file boundaries, see the sitemap partitioning diagnostic system. Those are adjacent operational topics; the specific method here is the provenance trace back to the origin-construction layer.
Separate remediation from detection
Do not change the proxy, application base URL and cache at the same time. Doing so can remove the symptom while making the responsible layer impossible to identify.
After the emitting layer is established, the appropriate correction may involve:
- setting the production public origin in the deployment environment;
- correcting an application or sitemap-specific base URL;
- removing a hard-coded or obsolete generator fallback;
- correcting trusted-proxy or forwarded-scheme handling;
- regenerating a static or scheduled sitemap artefact; or
- purging or replacing a stale cache and confirming that the new object is served publicly.
Choose the smallest change that fixes the responsible layer, then rerun the same evidence model. A redirect from an unintended hostname to production may make a URL retrievable, but it is not equivalent to direct emission when the production contract requires the sitemap to list the intended origin. This is an operational validation standard, not a claim that every redirecting sitemap URL causes search harm.
Define success as persistence, not a clean spot check
The correction is successful when:
- the intended origin is emitted consistently in sitemap indexes and all partitions;
- no unintended host, protocol or port remains in the complete sitemap graph;
- listed resources are retrievable as expected and their redirect behaviour matches the architecture;
- the fresh build or generation job produces the corrected values;
- the live public request path serves the corrected representation after cache handling; and
- the validation continues to pass on the next build, regeneration and deployment.
This standard deliberately goes beyond checking whether one URL now looks right. It tests whether the value is correct at source, remains correct in the generated artefact and survives delivery to the public sitemap endpoint.
What hostname drift establishes, and what it does not
Sitemap URLs are useful discovery hints, but the evidence does not support a universal claim that a wrong-host sitemap causes ranking loss, deindexation or crawl failure. Search engines may discover pages through other links and may use canonical, redirect and response signals when selecting URLs. Google’s sitemap guidance and duplicate-URL documentation support caution about treating sitemap inclusion as a decisive canonical signal.
That does not make hostname drift harmless as an operational defect. It can make the sitemap describe the wrong public architecture, obscure which URLs a search engine is being asked to crawl and expose a mismatch between deployment configuration and published output. The measurable SEO consequence must be investigated for the specific site rather than assumed from the hostname pattern.
The practical distinction is the important one: detection tells you that an origin is wrong; provenance tells you which layer supplied or preserved it. Once that layer is known, remediation can be narrow, and production validation can prove that the correction persists across builds, partitions, caches and live requests.
Share this article