Indexing lag after a release: how to tell delay from a discoverability defect

A release can leave new or changed URLs absent from search for several different reasons. This methodology uses release cohorts, control URLs and time-sequenced evidence to distinguish normal processing delay from a discoverability or indexability defect.

A new set of URLs is live, the release passed its deployment checks and yet the pages are not appearing in search. The immediate response is often to call this an indexing problem and request indexing for every URL.

That diagnosis is too broad to guide a useful response. The same symptom can result from normal processing delay, weak discovery signals, blocked crawling, a noindex directive, canonical selection, a template defect or an incomplete deployment. A successful 200 response does not settle the question.

This article sets out a narrower post-release method. Treat indexing diagnosis as a time-sequenced, cohort-controlled investigation: define the affected URLs, compare them with suitable controls and establish which observable transition has stalled: discovery, crawling, processing and indexability, canonical selection or search visibility.

The aim is not to estimate Google’s internal processing queue. That is not observable from the available tools. The aim is to decide whether the evidence supports waiting, correcting an implementation problem, escalating an investigation or validating a fix.

Start with a release cohort, not an isolated URL

Define the affected URL set using a clear inclusion rule. For example, a synthetic ecommerce release might include:

  • all 240 product URLs created or materially changed in the 14:00 deployment;
  • the product-detail template and its associated structured-data changes;
  • the XML sitemap update published at 14:12; and
  • the first category and related-product links exposed at 14:25.

This is the release cohort. Record the release identifier, deployment time, URL pattern, template version and intended indexability. Do not mix deliberately non-indexable pages, such as account, checkout or temporary transaction URLs, with pages intended to enter search.

Next, select unaffected control URLs. These might be comparable product pages that were live before the release, use the same template and remain in similar category positions. Controls should be comparable in:

  • template and content type;
  • internal-link context and approximate click depth;
  • sitemap treatment;
  • indexability rules; and
  • normal traffic or publication frequency, where those differences are material.

This control group is a methodological choice rather than a Google requirement. It is useful because the release cohort tells you what changed, while the controls provide a reference for the site’s usual crawl and visibility pattern. A small site may need to use all suitable comparable URLs. A large site can sample enough URLs to identify a pattern without treating any universal sample size as authoritative.

For general background on how crawling and indexing fit together, see our overview of how Google crawls and indexes a site. The method here is deliberately narrower: it starts with a known release and follows the available evidence forward.

Keep five different questions separate

“Is the page indexed?” compresses several different questions into one. Separate them before investigating:

  1. Discovery: has Google been given a route through which it could learn that the URL exists?
  2. Crawling: has Googlebot requested the URL?
  3. Processing and indexability: what did the crawler receive, and were there directives or access problems?
  4. Canonical selection: which URL does Google regard as the representative version of substantially similar content?
  5. Index visibility: is the URL, or a selected equivalent, currently visible in the search systems being checked?

These stages overlap operationally but are not interchangeable. An internal link and a sitemap entry expose a URL as potential discovery inputs; they do not prove that Google has discovered or crawled it. A verified Googlebot request proves that a request occurred and shows the response received; it does not prove rendering completion, canonical selection or index inclusion. A live URL Inspection test shows the current fetch outcome; it does not guarantee that the page will enter the index.

The useful question is therefore not simply whether the URL is absent from search. It is which evidence gate has not been passed, or whether the available evidence is still insufficient to identify the stalled transition.

Build a time-sequenced evidence chain

Record timestamps in one release log rather than investigating each source in isolation. At minimum, capture:

  1. deployment completion;
  2. the time the changed URL became publicly available;
  3. the first time an internal-link pathway was present on a crawlable page;
  4. XML sitemap publication or update;
  5. verified Googlebot requests and the responses returned;
  6. URL Inspection observations, separating the indexed version from the live test;
  7. declared canonical signals and any reported Google-selected canonical; and
  8. later index or search-visibility signals.

The order matters. If the page was absent from navigation until after the sitemap was generated, the sitemap may be the first available discovery input. If the URL was linked and included in a sitemap but has no verified crawler request, the evidence points towards an unresolved discovery or scheduling question, not automatically towards a processing failure. If Googlebot requested the page and received a valid response, move beyond the basic question of whether the URL can be fetched.

Confirm that the release reached production

Start with deployment and public-availability evidence. Fetch a sample of cohort URLs from outside the deployment environment and check that the expected template, content and directives are present.

Compare at least one affected URL with an equivalent control. A release can be marked successful by the deployment system while a production edge, cache, origin or feature flag still serves an older or incomplete version to some requesters. Look for:

  • the expected page content and template version;
  • a stable HTTP response rather than intermittent errors or redirects;
  • the intended canonical element;
  • the intended robots directives, including the absence of an accidental noindex;
  • the expected internal links; and
  • the correct sitemap membership.

A 2xx response, including 200 OK, is useful evidence that the URL was served successfully. It is not evidence that the URL will be indexed. Indexing can still be affected by directives, duplication, canonicalisation, content processing and other systems.

Test the discovery pathways

For each sample URL, record where Google could have encountered it. The main signals in this methodology are internal links and XML sitemaps.

A standard crawlable internal link generally uses an HTML anchor with an href attribute. Check the served and, where relevant, rendered output, and identify the linking pages rather than recording only that “the site links to it”. Were those linking pages live and accessible? Were they themselves linked from established sections of the site? Did the link appear before or after the sitemap update?

An XML sitemap is an additional way to communicate URLs and can provide a canonical hint, but inclusion is not a guarantee that Google has discovered, crawled or indexed the URL. Log sitemap presence as a discovery input, not as proof of progress.

These combinations help organise the investigation:

  • Internal link present, sitemap present, crawler request observed: the URL has passed several observable discovery gates. Investigate the served response, directives, canonical signals and subsequent processing.
  • Sitemap present, no meaningful internal link, no crawler request: discovery has been offered through one channel, but there is no URL-level evidence that Google acted on it. Compare with controls before deciding whether this is normal delay or a release-specific weakness.
  • Neither internal link nor sitemap exposure: this is a strong implementation candidate. Correct the missing pathway before treating the absence from search as a processing problem.

None of these combinations reveals Google’s internal scheduling. They organise the evidence that can be observed.

Establish crawler activity from logs

Server logs are the strongest source in this sequence for URL-specific crawl activity, provided they cover the relevant host, time period and infrastructure layer. Search Console Crawl Stats can add aggregate context about crawl activity and response categories, but it cannot demonstrate that a particular URL in the release cohort was crawled.

For sampled cohort and control URLs, record:

  • request time and URL;
  • user agent and IP information;
  • HTTP status and redirect chain;
  • response size or other useful delivery indicators; and
  • the infrastructure layer that recorded the request.

Do not accept a user-agent string alone as proof of Googlebot. Where the distinction matters, verify the request using reverse DNS followed by forward confirmation, or the relevant published IP information. Also confirm that logs include CDN, edge and origin requests consistently. A missing origin entry may mean that the edge served a cached response, not that Googlebot did not request the URL.

A verified Googlebot request establishes that Google’s crawler requested the URL at a particular time and received a particular response. It does not establish that Google rendered the page, selected it as canonical or added it to the index. If the release cohort has requests comparable with controls, the issue is less likely to be a simple absence of crawl opportunity, although other defects can still exist.

Use URL Inspection without treating it as an oracle

URL Inspection contains two different observations that should be recorded separately:

  • The indexed version: what Search Console reports about the version Google previously processed. This can be stale relative to the current production page.
  • The live test: what the current test fetch can access and observe. This reflects the present state, but does not guarantee inclusion or predict Google’s canonical choice.

It is possible for the indexed version and live test to disagree without either result being defective. A current live test may show that the page can be fetched while the indexed report still reflects an earlier response, an older directive or no indexed version.

Use the live test to identify current access, response and directive problems. Use the indexed report to understand the last reported indexed state. Do not turn a result such as “URL can be indexed” into a forecast that it will be indexed next, and do not treat a failed live test as proof that the original release failed in the same way unless the timeline supports that conclusion.

Inspect canonical signals as a separate decision point

A URL can be technically accessible and still not become the representative indexed URL. Compare:

  • the declared canonical;
  • the canonical signals implied by internal links and sitemap inclusion;
  • the URL’s substantive content compared with controls and nearby variants; and
  • any Google-selected canonical reported in URL Inspection.

Google may select a different representative URL from the one declared by the site. A mismatch does not, by itself, identify whether the cause is duplication, conflicting signals, content similarity or another processing decision. Treat it as evidence that canonical selection needs investigation, not as proof of a single technical fault.

If Googlebot has crawled the URL and received the intended content, continued non-inclusion should be investigated as a possible canonical, duplicate, content or processing outcome rather than automatically labelled a crawl defect. The available tools may not expose the exact reason, so compare the cohort with equivalent templates and examine whether a different URL is receiving the expected visibility.

Choose the response from the evidence

There is no defensible universal waiting period that applies to every release. Crawl and processing patterns vary by site, URL type, update frequency, infrastructure health and other conditions. Use controls and the site’s own history instead of turning an arbitrary number of hours or days into a rule.

The response should fall into one of four categories: wait, correct, escalate or validate.

Wait and recheck when the evidence is incomplete but healthy

Waiting is reasonable when:

  • the release is confirmed in production;
  • the URL has an intended internal-link or sitemap pathway;
  • there is no deterministic block such as noindex, a robots exclusion, an error response or a broken redirect;
  • the release is recent relative to comparable site activity; and
  • the control group shows similar processing lag, or the site has insufficient historical evidence to identify a material deviation.

Set a specific recheck point based on the site’s normal release pattern. “Wait” should mean continuing to observe the defined evidence chain, not closing the issue.

Correct a deterministic defect

Correct the implementation when the evidence identifies a reproducible problem, such as:

  • the cohort is missing from the intended sitemap;
  • new URLs are not exposed through the expected internal-link pathway;
  • the production template returns an accidental noindex;
  • robots rules prevent the relevant fetch;
  • the URL returns an error, redirect or incomplete content unexpectedly;
  • the canonical points to the wrong variant; or
  • the deployment reached only part of the intended URL set or served inconsistent template versions.

Fix the cause first. A manual request for indexing cannot compensate for a missing pathway, blocked crawl or incorrect directive.

Escalate a material deviation from controls

Escalate to engineering, product or a senior SEO owner when the release cohort shows a pattern that controls do not. Examples include:

  • equivalent control URLs continue to receive crawler requests while the entire cohort does not;
  • only URLs using the changed template return a different status, directive or canonical;
  • the cohort has discovery signals but a sustained lack of crawl activity relative to comparable releases; or
  • live tests look correct, but the indexed state remains inconsistent across the cohort after the site’s normal reprocessing pattern.

Escalation should include the cohort definition, control selection, timeline, representative URLs, log evidence, rendered and served checks, URL Inspection observations and the specific decision that remains unresolved. Avoid escalating with only a list of URLs labelled “not indexed”.

Validate after the fix

After correcting an issue, rerun the same checks against the affected cohort and a control sample. A successful live test confirms that the current fetch has improved; it does not prove that the indexed version has already changed. Record the next observable crawler request, the response received, canonical observations and later visibility signals.

Requesting indexing through Search Console can be appropriate for a small number of important URLs after the implementation is correct. Treat it as a recrawl request, not as proof that the defect is fixed or that inclusion is guaranteed.

Shortcuts that produce weak diagnoses

“It is not in a site: search, so it is not indexed.”

Search operators can provide a supporting visibility signal, but they are not a complete diagnostic. Absence from a site: search is not conclusive, and presence does not explain how the URL was discovered, whether it is canonical or whether the current version is represented.

“It is in the sitemap, so Google knows about it.”

A sitemap entry tells Google that the publisher is communicating the URL. It does not prove that Google has discovered, crawled or indexed it. Treat sitemap publication as a timestamp in the release timeline.

“The live test passed, so indexing is imminent.”

A live test that can fetch the page is useful evidence against some current access and directive problems. It does not evaluate every indexing condition, guarantee inclusion or predict canonical selection.

“Googlebot visited it, so the page should be visible.”

A verified request is an important milestone, but crawling is not the same as processing, canonical selection or index visibility. Continue the investigation rather than treating the log entry as the final answer.

Prepare the next release

  • Define the release cohort before deployment, including its URL pattern, template and intended indexability.
  • Select comparable control URLs or templates and record why they are comparable.
  • Save pre-release samples of served HTML, rendered output, status, directives, canonical and internal links.
  • Confirm which internal-link pathways should expose the new or changed URLs and when those pathways become public.
  • Record sitemap generation, publication and inclusion times.
  • Confirm that logs cover the relevant CDN, edge and origin layers and that crawler verification is possible.
  • Capture baseline crawl and visibility patterns for the controls and recent comparable releases.
  • After release, record verified crawler requests and responses for a sample of cohort and control URLs.
  • Run URL Inspection selectively, keeping indexed-version observations separate from live-test results.
  • Compare declared and selected canonical signals without assuming that a mismatch has one known cause.
  • Classify the outcome as wait, correct, escalate or validate, with the evidence supporting that decision.
  • After any fix, repeat the same checks and record what changed before judging the release outcome.

Conclusion

“Not indexed yet” is an observation, not a diagnosis. The useful distinction is whether a release is still moving through normal, partly observable processing or whether a specific implementation defect has removed or weakened a required pathway.

Define the cohort, compare it with controls and build the timeline from production availability through discovery signals, crawler activity, served responses, URL Inspection, canonical selection and later visibility. That sequence will not reveal Google’s private processing decisions, but it can show where the observable chain stops and whether the next action should be to wait, correct the release, escalate the investigation or validate a fix.

Share this article

Found this useful? Pass it on.

Share on LinkedIn · Share on X