Vary Headers and SEO: Diagnosing Cache-Dependent HTML Differences

A practical method for testing whether HTTP negotiation and intermediary caches serve materially different HTML, metadata or structured data for the same URL.

A page can appear correct in one test and still return different HTML or metadata to another requester. Possible causes include language negotiation, device adaptation, an experiment, a stale edge object, bot-specific application logic or a cache key that does not match the representations being served.

This creates a specific diagnostic problem. The question is not simply whether two browsers display different pages. It is whether the same URL returns a materially different HTTP response for a defined request profile and cache state, and whether that difference changes an SEO-significant signal.

The method below treats the request, response, delivery path and cache state as one unit of analysis. It uses repeated requests, response comparison and edge or origin evidence to separate competing explanations.

Analyse the effective HTTP response

For this investigation, a URL is not the complete object being tested. The unit of analysis is:

  • the exact URL and HTTP method;
  • the request headers, cookies and other request characteristics;
  • the delivery path, such as CDN, reverse proxy and origin;
  • the cache state at the time of the request; and
  • the effective HTTP response received.

That last point matters because a browser display is the result of more than the initial response. JavaScript, service workers, local storage and browser caching may alter what is eventually visible. Conversely, two raw responses may differ in whitespace, compression or dynamic identifiers while producing the same SEO-relevant page.

Start with the raw response, not the rendered appearance. Record the request that produced it and preserve the response headers and body. A browser can be useful later, but it should not be the first or only diagnostic instrument.

What Vary tells you

HTTP permits a server to provide different representations of a resource in response to request characteristics. Common examples include Accept, Accept-Language and Accept-Encoding. The HTTP specification describes the semantics of Vary in its section on content negotiation: RFC 9110, section 12.5.5.

The Vary response header identifies request fields that may have influenced representation selection. It gives a cache information about which request fields to consider when deciding whether a stored response is suitable for a later request.

HTTP/1.1 200 OK
Vary: Accept-Language, User-Agent
Cache-Control: public, max-age=300

Treat Vary as a cache-matching instruction, not as proof that every observed response difference is valid or correctly isolated. It does not prove that:

  • the declared dimensions are the only inputs affecting the response;
  • the application has selected the right representation;
  • the CDN uses the same cache-key logic as the origin; or
  • the resulting variants differ in a way that matters for SEO.

A missing Vary header does not prove that a response is invariant either. An application, reverse proxy or CDN may vary on cookies, geography, experiments, device signals or internal state. The header is evidence about cache matching, not a complete description of the delivery system.

Vary: * has a more restrictive meaning: other aspects of the request may have influenced the response, so a cache cannot determine that a stored response is suitable for a later request without forwarding that request to the origin. Operational behaviour should still be verified rather than assumed from the header alone.

CDNs and reverse proxies may apply provider-specific rules to headers, cookies, query strings, geography, device signals and other inputs. Those rules can add to, exclude or normalise the dimensions suggested by the origin response. The effective cache key therefore needs to be investigated separately.

Keep the possible causes separate

Do not begin by assuming that the CDN is at fault. The same visible difference can be produced or preserved by several layers:

  • Origin variation: application logic selects different HTML for a language, device, user-agent, cookie, location or experiment.
  • Reverse-proxy variation: an internal proxy applies its own routing, caching or normalisation rules.
  • CDN variation: an edge cache stores or retrieves representations using a cache key that differs from the origin’s assumptions.
  • Stale content: one cache tier continues to serve an older object after a deployment or content update.
  • Client variation: a browser cache, service worker or local state changes what the user sees after the response has arrived.
  • Rendering variation: the initial HTML is equivalent, but client-side code produces a different final page.
  • Legitimate negotiation: language, format or device-specific representations are intentional and correctly isolated.

Each hypothesis predicts different evidence. A single browser request, or a single response containing Vary, cannot settle the diagnosis.

Build a controlled request matrix

Vary one dimension at a time before testing combinations. The relevant dimensions depend on the site, but a useful starting point includes:

  • the normal desktop user-agent;
  • a mobile user-agent;
  • a search crawler user-agent, as a simulation only;
  • Accept values relevant to the application;
  • Accept-Encoding, such as gzip and Brotli-capable requests;
  • Accept-Language values used by the site;
  • device or client-hint signals where the site uses them;
  • anonymous and relevant cookie states;
  • requests with and without cache-bypass controls, where those controls are supported;
  • different source regions or edge locations, where regional delivery is relevant; and
  • repeated requests at defined intervals.

Keep the URL, method and other variables fixed while testing a single dimension. Then test combinations that reflect real traffic. For example, a language difference may only appear when a particular user-agent and cookie are present.

A command-line capture makes the request explicit:

curl --compressed -sS -D response.headers \
  -H 'User-Agent: ExampleBrowser/1.0' \
  -H 'Accept-Language: en-GB' \
  -H 'Accept-Encoding: gzip, br' \
  'https://www.example.com/example-page' \
  -o response.html

For a crawler profile, use the relevant user-agent string only as a test input. A command-line request that claims to be Googlebot is a simulation, not proof that verified Googlebot traffic receives the same response. Google explains that user-agent strings can be spoofed and that crawler verification requires additional checks: Google’s crawler-verification guidance.

For each request, save the timestamp, URL, request headers, source location, cookies, response headers, redirect chain, body and any correlation ID. Avoid relying on a tool that silently follows redirects or discards headers without recording those steps.

Compare transport evidence with page meaning

Comparison should move from transport-level evidence to semantic HTML. Start with:

  • status code;
  • redirect chain and final URL;
  • content type and content encoding;
  • Cache-Control, Vary, ETag, Last-Modified, Date and Age;
  • Cache-Status, Via and vendor-specific cache fields where exposed; and
  • request or response correlation IDs.

Cache-Status is a standardised response field for reporting cache handling. Its syntax and intended use are defined in RFC 9211. Age, ETag, Date, Via and provider-specific fields can add useful evidence. None is universal proof of which layer supplied the response: headers may be absent, rewritten or exposed only by selected parts of the delivery path.

Compare both the compressed and decompressed body where practical. A raw byte hash can show that something changed, but it cannot explain whether the change is meaningful. Compression, whitespace, timestamps, nonces and dynamic identifiers may produce different hashes without changing the page’s SEO function.

Use a normalised HTML comparison to reduce that noise, then extract specific fields:

  • the title and meta description;
  • robots directives in the HTML and relevant HTTP headers;
  • canonical URL;
  • hreflang annotations;
  • structured data types, entities and key properties;
  • primary heading and other primary content;
  • indexable text;
  • internal links and their destinations; and
  • links or elements that control discovery of important resources.

This field-level comparison answers a more useful question than “are the files identical?” It shows whether the effective representation changes indexability, canonicalisation, discovery, primary content or search-feature inputs.

Test cache state directly

Cache state is central to the investigation. Test a sequence rather than taking one snapshot.

  1. Capture a baseline. Make repeated requests with the same profile and record whether the response is stable.
  2. Test cold and warm behaviour. Where safe, use a controlled purge or a dedicated test object. Make the first request and repeat it from the same path. Compare the body, headers, age and cache indicators.
  3. Test another profile. Change one request dimension, such as language or user-agent, and repeat the cold-to-warm sequence.
  4. Test revalidation. Use the validators returned by the server, such as ETag or Last-Modified, and record whether the intermediary returns a fresh representation, a 304 response or an older object.
  5. Repeat over time and across edges. Different points of presence or cache tiers may hold objects with different ages. A MISS followed by a HIT may reflect distributed cache state rather than a new origin response.

Interpret freshness, revalidation and stale-serving behaviour against the active cache directives and intermediary configuration. An older response does not, by itself, prove that the cache key is incorrect.

Purging needs careful interpretation. A purge confirmation may apply to one region, tier or key and may not prove global invalidation. Record the edge or region involved and allow for propagation where appropriate.

Cache-bypass headers are not automatically authoritative. A request header may be ignored, normalised or handled differently by each intermediary. Treat it as an experimental condition, then verify the result through response headers and logs.

Where infrastructure access exists, include a unique request correlation ID and compare public responses with reverse-proxy and origin logs. A useful sequence is:

  • send a request with a unique ID;
  • confirm whether that ID reached the origin;
  • repeat the request and see whether the second response is served without reaching the origin;
  • compare the response body returned from the public path with the origin response; and
  • repeat the same test for each request profile that produced a difference.

If one variant reaches the origin and another is served from an edge without an origin request, that is useful evidence. It is not, by itself, proof of a faulty cache key. The origin may intentionally generate the first representation and the cache may be correctly serving it.

Use a control URL and repeat the samples

A control reduces the risk of mistaking a transient change for cache-dependent variation. Choose a URL that should be stable and is delivered through the same relevant infrastructure. Run it through the same request matrix and timing sequence.

For example, if the target product page changes only for a mobile user-agent after a purge, but the control page shows the same unexplained pattern, the issue may be broader edge state or an infrastructure change. If only the target URL differs and the origin logs show an application experiment, the cache is less likely to be the primary cause.

Repeat samples before and after deployments where possible. A single before-and-after comparison cannot distinguish a stale object from an editorial update, A/B test, personalisation or a deployment that changed the origin output.

Record evidence against each explanation rather than assigning a cause from the first plausible clue:

  • Application logic: the origin returns different HTML for matched request profiles, and the difference is visible on a direct or otherwise controlled origin path.
  • Language negotiation: changing Accept-Language consistently changes the representation, with the relevant dimension declared or deliberately handled.
  • A/B testing or personalisation: a cookie, experiment assignment or user state predicts the difference, and the origin or application logs support it.
  • CDN or proxy behaviour: the public path returns different content while matched origin responses are stable, with cache indicators or edge logs showing separate retrieval behaviour.
  • Stale objects: an older body, age or validator persists after the origin changes, and the object changes after the relevant purge or expiry event.
  • Rendering: raw responses are materially equivalent, but the browser or JavaScript execution produces different final output.
  • Bot handling: the application branches on the supplied user-agent, although a simulated crawler request does not prove the behaviour of verified crawler traffic.

Judge the SEO consequence

Not every response difference is an SEO defect. Different compression, cache metadata, whitespace or dynamic identifiers may have no meaningful search consequence.

The relevant question is whether the response changes a signal that affects how the page can be understood, indexed or discovered. Prioritise differences in:

  • status code or redirect destination;
  • robots directives;
  • canonical target;
  • hreflang relationships;
  • title and other important metadata;
  • primary content and headings;
  • internal links and crawl paths; and
  • structured data that describes a materially different entity, product, offer or content type.

A cache-dependent difference deserves more attention when it changes one of these fields for a request profile that a search engine may use, or when it creates inconsistent representations across repeated public requests. A confirmed delivery difference still does not prove a ranking, indexing or traffic outcome. It establishes a delivery condition that needs to be assessed alongside crawler evidence and search-engine processing.

For example, a product page that returns the same product information but different Brotli and gzip byte sequences has a transport difference, not necessarily an SEO issue. A page that returns noindex, a different canonical or a materially incomplete product description to one request profile requires a more urgent investigation.

For related edge-routing failure modes, see how CDN edge traces can expose redirect loops outside application code. If the suspected cause is regional routing rather than cache-dependent representation selection, see the separate method for diagnosing geo-IP redirects.

Turn the finding into an implementation test

Once the evidence identifies the likely layer, the fix should reflect the representations the site intends to serve. There is no universal cache policy.

Review:

  • which request dimensions genuinely select a different representation;
  • whether those dimensions are declared through Vary where required;
  • whether the CDN and reverse-proxy cache keys include the same effective dimensions;
  • whether high-cardinality headers, cookies or client hints are fragmenting the cache unnecessarily;
  • whether private or personalised responses can enter a shared cache;
  • whether freshness, revalidation and stale-serving rules match the publishing workflow;
  • whether purge operations cover the relevant keys, regions and cache tiers; and
  • whether the origin, edge and monitoring systems expose enough information to diagnose a future mismatch.

Regression testing should use the same request matrix that found the issue. For each important URL type, assert the expected status, redirect chain, cache policy and SEO-significant fields for the supported profiles. Include cold and warm sequences, repeated requests and at least one control URL. Where variation is intentional, test that the correct variants remain isolated rather than requiring every response to be byte-for-byte identical.

Monitoring should alert on changes that matter: unexpected canonical or robots values, missing structured data, unusual body signatures, inconsistent primary content or a cache hit serving a representation outside its intended profile. A raw hash can help identify a change, but field-level checks are needed to decide whether it is important.

Conclusion

The key distinction is between a different response and a different browser display. Investigate the former by pairing the exact request profile with cache state, delivery path and effective HTTP response.

Vary is useful evidence because it describes cache-matching inputs, but it is not a guarantee that the cache key is complete or that a response difference is legitimate. The strongest diagnosis comes from triangulation: controlled requests, cold and warm sequences, response and field-level comparison, control URLs and origin or edge evidence where available.

Only after establishing what changed should you assess SEO impact. The practical objective is not to eliminate all variation. It is to ensure that intentional representations are correctly isolated, stale or colliding objects are detected and every request profile receives the HTML and metadata that the site intends to publish.

Share this article

Found this useful? Pass it on.

Share on LinkedIn · Share on X