Structured data @id drift: keeping entity identifiers stable across releases

A practical guide to designing deterministic JSON-LD @id values, diagnosing graph fragmentation and testing entity continuity across templates, environments and releases.

A structured-data deployment can remain syntactically valid while quietly changing the identity of the entities it describes. An @id built from the request hostname, locale path, preview flag or deployment version may identify a different JSON-LD node after a release, even when the visible content appears unchanged.

That matters when reusable entities such as an organisation, website, author, product or article are described in several templates and referenced by more than one graph. If their identifiers drift, the publisher’s graph becomes harder to reconcile and govern. A practical response is to treat reusable @id values as versioned data contracts: define them centrally, keep them independent of request context and test them as part of release engineering.

This guide explains what an @id does, how it differs from a page URL and visible name, how to design durable identifiers for five common entity types, and how to validate continuity across source HTML, rendered output, production responses, locales and deployments.

What @id means in a JSON-LD graph

In JSON-LD, @id identifies a node using an IRI reference, compact IRI or blank-node identifier. An IRI is the identifier form used by linked-data systems. It may look like a URL, but its role is to identify a node in the graph rather than simply provide a destination for a browser.

The JSON-LD 1.1 specification also defines how relative identifiers are resolved against a base IRI. An identifier such as /entity/acme can therefore resolve differently when a document is processed at https://www.example.co.uk/about and at https://preview.example.net/about. The strings look identical in the source, but the resulting absolute identifiers differ.

Consider this node:

{
  "@type": "Organization",
  "@id": "/entity/acme",
  "name": "Acme Travel"
}

On the production site, it may resolve to https://www.example.co.uk/entity/acme. In a preview environment, it may resolve to https://preview.example.net/entity/acme. If both documents are intended to describe the same organisation, the identifier has become environment-dependent.

An explicit, absolute identifier avoids that particular ambiguity:

{
  "@type": "Organization",
  "@id": "https://id.example.co.uk/organization/acme",
  "name": "Acme Travel"
}

This does not mean every site must use an /organization/ namespace. A governed canonical page URL can also be a valid identifier when an entity is deliberately tied to that page. The design question is whether the identifier should survive changes to the page URL, slug, locale or template. The RDF 1.1 Concepts specification describes the persistence principle behind this choice: an IRI intended to identify a resource should not change its intended referent.

Keep identity separate from related signals

Four values are commonly confused during implementation:

  • @id: the identifier of a node in the JSON-LD graph.
  • url: a property describing a URL associated with that entity, often its public page.
  • Canonical URL: the URL signal a page supplies to search engines as its preferred version.
  • Visible name: the human-readable label, such as “Acme Travel” or “The Mountain Guide”.

A single URL can deliberately serve more than one role. For example, an article may use its canonical page URL as both url and @id if its identity is page-bound and URL governance is strong. Those roles should still be documented separately. A URL migration should not silently change the entity identifier unless that is an intentional identity change.

Names are particularly poor identifiers. They can change for editorial, legal or localisation reasons, and two different entities can share the same name. A product title such as “Trail Pro Jacket” is not a sufficient key for a product record. A visible name is evidence about an entity; it is not necessarily the entity’s durable identifier.

Similarly, @id is not a ranking signal to optimise in isolation. Google’s structured-data policies explain that structured data can make a page eligible for certain search appearances, but does not guarantee rankings or a rich result. Stable identifiers are best understood as a publisher-controlled way to express continuity and relationships, not as a guarantee of search visibility.

Design rules for deterministic identifiers

A deterministic identifier returns the same value whenever the same logical entity is requested, regardless of the page template or deployment context. Achieving that requires an explicit identity policy rather than string concatenation inside individual templates.

1. Choose a deliberate namespace

Use a publisher-controlled namespace that makes ownership and entity type clear. For example:

https://id.example.co.uk/org/acme-travel
https://id.example.co.uk/site/main
https://id.example.co.uk/person/maya-patel
https://id.example.co.uk/product/sku-84721
https://id.example.co.uk/article/guide-to-alpine-rail

Separate namespaces can reduce accidental collisions and make audits easier, but this is an implementation convention rather than JSON-LD or Google-mandated syntax. The useful properties are that the namespace is deliberate, the key is stable and the ownership rules are documented.

2. Base the key on a stable entity record

Prefer an immutable database identifier, controlled slug or assigned registry key. Do not derive the key from a mutable title, request path or arbitrary array position.

Good:    https://id.example.co.uk/product/84721
Risky:   https://id.example.co.uk/product/trail-pro-jacket
Bad:     https://www.example.co.uk/en-gb/jackets/trail-pro-jacket?build=2025-03-14

A slug may be acceptable where slug changes are governed and aliases are maintained. An internal record key is generally easier to preserve through renames and URL migrations.

3. Exclude deployment and request context

Unless the data model intentionally defines separate entities, exclude hostname, preview mode, build hash, deployment ID, template path, query parameters and incidental locale segments from the durable identifier.

These values should not identify the same organisation:

https://preview.example.net/entity/acme?preview=true
https://www.example.co.uk/entity/acme?deploy=9f2a31

Both environments should instead emit the same approved organisation identifier, while their page URLs and delivery endpoints may differ.

4. Decide where context creates a genuinely different entity

Environment independence does not mean every entity across every market must share one identifier. Separate regional organisations, websites, translated works or product variants may be legitimate distinct entities. The distinction needs to be modelled deliberately.

For example, a global organisation and its legally separate French subsidiary should not be forced into one node merely because their names are similar. Conversely, a translated version of the same editorial article should not receive a new identity simply because the URL contains /fr/ if the publisher’s model treats it as one content record with multiple language representations.

Document the decision in the entity registry. Do not let the URL structure make it by accident.

5. Treat blank nodes as local structures, not reusable identities

Blank nodes are valid JSON-LD and can be useful for anonymous structures that do not need cross-document references. They are not a substitute for a durable, reusable IRI. The JSON-LD specification distinguishes blank-node identifiers from named IRI-based nodes, so a local anonymous object should not be expected to provide a stable cross-template identity.

Worked identity cases

The following examples focus on identifier lifecycle rather than the full property requirements for each schema.org type.

Organization

An organisation normally needs one documented primary identifier wherever the same organisation is referenced:

"@id": "https://id.example.co.uk/org/acme-travel"

A common defect is to generate the value from the current site host:

"@id": "https://{{ request.hostname }}/organization"

This creates separate nodes for production, staging and regional hosts. If the business model intentionally treats each regional legal entity as separate, use separate registry records. If not, the hostname should not be part of the identifier.

WebSite

A WebSite identifier should represent the website entity defined by the publisher’s model, not whichever hostname served the request:

"@id": "https://id.example.co.uk/site/main",
"url": "https://www.example.co.uk/"

The url may change during a domain migration while the website record remains the same, provided continuity is intended and the migration is governed. A multilingual site may instead define regional websites separately. Either model can be valid; the release system needs to know which one applies.

Person

For a reusable author or contributor record, avoid creating an ID from the display name:

Unstable: https://id.example.co.uk/person/maya-patel?locale=fr
Stable:   https://id.example.co.uk/person/1842

The stable key allows the person’s name, profile URL or biography to change without changing the node identity. It does not, by itself, prove that two similarly named people are the same person. That remains an editorial and data-governance decision.

Product

Product identity requires particular care because a product family, a sellable variant and an offer may be different records. Use a controlled product or variant key according to the commerce model:

Product family: https://id.example.co.uk/product/trail-pro-jacket
Variant:        https://id.example.co.uk/product/trail-pro-jacket/blue-medium

Do not let a currency, stock state or campaign parameter create a new product identity unless the business has explicitly defined a separate entity. Price and availability can change while the product identifier remains stable. Google’s Product structured-data documentation describes the product information to represent, but does not prescribe one universal @id pattern or establish that a particular identifier design improves rankings.

Article

For an article, decide whether the identifier represents an editorial content record, a language-specific work or a page-bound publication:

Content record: https://id.example.co.uk/article/rail-travel-guide
French work:    https://id.example.co.uk/article/rail-travel-guide/fr

Both choices can be defensible. The defect is allowing the framework to choose based on the current template path, such as /en/, /fr/ or a preview route, without a documented content model. Article dates, headlines and page URLs can change while the intended article node remains the same. The Article documentation is useful for feature-specific markup, but it does not define a universal lifecycle policy for article identifiers.

How identifier drift fragments a graph

Graph fragmentation occurs when one logical entity is represented by multiple active identifiers without an intentional reason. It can also occur when references point to identifiers that no longer match the node definitions emitted elsewhere.

For example, a homepage may describe:

{
  "@id": "https://id.example.co.uk/org/acme-travel",
  "@type": "Organization",
  "name": "Acme Travel"
}

While an article template references:

{
  "@type": "Article",
  "author": {
    "@id": "https://www.example.co.uk/about-us#organization"
  }
}

That is not automatically wrong. The two values might intentionally represent different nodes, or the second might be a legacy identifier with an explicit reconciliation policy. If both are meant to represent the same organisation, though, the graph has lost continuity.

Do not infer sameness from matching names alone. Graph syntax cannot determine whether two IRIs identify the same real-world entity. Record the intended relationship in the registry and classify the difference as an approved distinction, an alias requiring migration or a defect.

A diagnostic method for finding drift

A validator that reports valid syntax or feature-oriented markup may still miss cross-template and release-level identity changes. Use a diagnostic inventory that records, for every parsed identifier:

  • the absolute @id;
  • the node types;
  • the page, template, locale and environment where it appeared;
  • the release or build version;
  • incoming references from other nodes;
  • outgoing properties and their values;
  • whether the identifier exists in the approved registry.

JSON-LD expansion resolves relative IRIs and expands terms and compact IRIs, making it a stronger basis for comparison than raw JSON string matching. The JSON-LD API specification documents this processing model.

Run these checks against representative pages and releases:

  • Duplicate identity check: flag an identifier when it is used for unrelated records or incompatible types.
  • Multiple-ID check: identify records that the registry says are one entity but which appear under multiple active identifiers.
  • Dangling-reference check: find references to IDs that are absent from the tested graph or approved cross-document registry.
  • Environment-token check: flag hostnames, preview markers, query parameters, build hashes and deployment IDs in durable IDs.
  • Template-consistency check: compare shared entities emitted by homepage, article, product, category and profile templates.
  • Relationship check: confirm that references such as publisher, author, brand or isPartOf point to the intended registry IDs.

RDF Dataset Canonicalization can reduce noise from serialization order and blank-node labelling when comparing graph snapshots. It cannot decide that two different IRIs represent the same entity. For smaller implementations, a targeted inventory of IDs, types and references may be a more practical CI baseline than full dataset canonicalisation.

Test three delivery layers

Structured data can be present in server-rendered source, inserted or altered during hydration, or affected by CDN and production configuration. Google documents the use of JavaScript-generated structured data in its guide to generating structured data with JavaScript. One inspection layer is therefore not always enough.

Source HTML

Fetch the initial response and extract every JSON-LD block. Check whether the server-side template emits the expected absolute IDs and whether environment values have leaked into them.

Rendered DOM

Load representative pages in a browser and inspect the post-hydration DOM. Compare the graph with source HTML to detect client-side duplication, replacement or identifier changes.

Production responses

Test the public response through the real CDN, redirects, cache variants and canonical host. A clean staging response does not prove that production configuration emits the same graph.

Some differences are expected. A preview page may have a different page URL, a locale may have different translated properties and a deployment may change dateModified. None is automatically an identity defect. A production organisation ID that changes because the request hostname changed is usually a defect when the organisation is intended to be shared.

Make @id a release data contract

The central implementation recommendation is to stop allowing templates to reconstruct reusable identifiers independently. Maintain a registry or equivalent source of truth containing the public ID, entity type, canonical relationships, lifecycle status, aliases and approved regional or language distinctions.

Then add four controls to the release process:

  1. Fixtures: create representative records for each reusable entity and important content model, including regional, translated and variant cases.
  2. Invariants: assert rules such as “the primary organisation ID does not vary by hostname” and “the product ID does not include price, currency or campaign parameters”.
  3. Graph snapshots: compare expanded or canonicalised graphs across representative templates and releases. Allow expected property changes without allowing unapproved identity changes.
  4. Approval for intentional changes: require an explicit migration note when an identifier must change because the underlying entity has merged, split, retired or been redefined.

Shape validation tools such as SHACL can test structural constraints including types, cardinality and permitted relationships. Shape validation checks the graph supplied to it; it does not know the publisher’s cross-release identity policy. Custom assertions, registry comparisons or application-level tests are therefore needed for invariants such as stable IDs across environments.

A useful release report should show changed identifiers, affected templates and the reason for each difference. Do not make raw JSON equality the only gate: property order, whitespace and expected mutable properties can change without changing identity.

Definition of done

Identifier work is complete when the implementation can demonstrate all of the following:

  • Identity rules are documented for reusable Organization, WebSite, Person, Product and Article entities.
  • A registry or equivalent source of truth owns public IDs and records intentional distinctions.
  • Durable IDs are independent of hostname, preview mode, deployment version, template path and irrelevant query parameters.
  • Shared references resolve to the same approved IDs across representative templates and locales.
  • Source HTML, rendered DOM and production responses have been checked where the architecture can change the graph.
  • Expanded or parsed graph comparisons distinguish identity changes from expected property changes.
  • Duplicate IDs, multiple IDs for one governed entity and dangling references are classified and resolved.
  • CI or release checks fail on unapproved identifier drift and produce an actionable diff.
  • Intentional identity changes have an owner, migration note and documented reason.

The fixture set should match the site’s risk. A small site may need a handful of templates and entities; a multinational ecommerce platform may need coverage for markets, variants, translated content, preview modes and URL migrations.

What stable identifiers cannot solve

A stable @id does not make inaccurate properties correct. It cannot resolve conflicting canonical or page signals, supply missing required properties, establish weak entity evidence or control how a search engine interprets the graph. Search engines may also reconcile entities using URLs, visible content, feeds, product identifiers, sameAs and other signals.

Nor is every identifier difference a defect. A regional organisation, product variant or translated work may properly have its own node. The governing question is whether the identifier reflects the intended entity model and remains consistent with that model across releases.

Structured-data validation therefore needs two separate questions: is this graph structurally valid, and does it preserve the publisher’s intended identity relationships? The first can be supported by schema and shape validation. The second requires a registry, fixtures, graph comparisons and release ownership.

Conclusion

The useful distinction is between a page being validly marked up and a publisher preserving entity continuity. An @id is a graph identifier, not merely another spelling of a page URL or visible name. If it is generated from unstable request or deployment context, the same intended real-world entity can appear as multiple nodes across the site.

Design IDs from a deliberate namespace and stable entity key. Document where locale or product context creates a genuinely different entity, and keep the generation logic out of individual templates. Then test the parsed graph across templates, environments and releases, with explicit controls for expected changes.

For a broader operational view, see Liquid Silver’s guidance on turning structured-data validation findings into action and using template fingerprinting to identify SEO deployment drift. The wider implementation context is covered in SEO engineering.

Share this article

Found this useful? Pass it on.

Share on LinkedIn · Share on X