Crawled but not indexed: how to investigate the real cause
A practical, evidence-led method for investigating pages Google has crawled but not indexed, from canonical choices and soft errors to genuine processing delays.
A page exists, Google has crawled it, and yet the URL does not appear in Google Search. Search Console reports “Crawled – currently not indexed”, leaving the marketing team with an awkward question: what, exactly, is wrong?
The honest answer is that the status does not tell you. It confirms a processing state, not a complete algorithmic diagnosis. The page may be duplicated, consolidated with another URL, behaving like a missing page, lacking a distinct reason to appear, or simply waiting for further processing.
That distinction matters. Treat every excluded URL as a technical defect and you may spend weeks adding copy, resubmitting sitemaps and requesting indexing for pages that do not need their own search result. Assume every case is harmless and you may miss a broken template affecting thousands of URLs.
This article sets out a practical investigation sequence. It starts with one question: does this URL deserve to be indexed as a separate result? From there, it uses the available evidence to decide whether to fix, consolidate, improve, intentionally exclude, or wait and monitor.
What “crawled but not indexed” tells you, and what it does not
Search engines handle URLs in stages. Discovery is becoming aware that a URL exists. Crawling is requesting and processing the URL. Indexing is evaluating and storing a page, or a representative version of it, for possible inclusion in search results. Google’s overview of how crawling and indexing work explains the broader process.
For this investigation, crawling has already happened. The question is no longer simply whether Google can find or fetch the URL. It is how Google has interpreted the page and its relationship to other URLs.
Google defines “Crawled – currently not indexed” as a URL that Google crawled but did not currently include in the index. It does not say that the page is permanently rejected or identify one specific cause. Nor does it prove that a particular technical change will make the URL eligible.
Think of the label as a signpost rather than a verdict. Search Console reports describe a processing state at a particular point in time, and different tools can reflect different points in that process. A live URL test can confirm some technical signals, but it cannot predict every indexing outcome, including all canonical decisions or whether a URL will move into the crawled-but-not-indexed state.
The sensible response is evidence triangulation: compare the status with the live page, competing URLs, site architecture and the page’s commercial purpose.
Start with the business question: should this page have its own result?
Before inspecting source code, decide what the URL is supposed to do.
A product detail page may deserve visibility because it represents a genuinely different product. A location page may be useful when it serves a distinct area with relevant information. A filtered ecommerce URL, temporary campaign page or near-identical service page may be useful to visitors who reach it through the site but still not warrant a separate organic result.
Google does not need to index every URL a site can generate. Its canonicalisation guidance explains that search engines often select a representative, or canonical, URL from duplicate and similar versions.
Write down the intended outcome before making changes:
- Separate visibility: the page answers a distinct search need and should be eligible for its own result.
- Consolidation: another URL is the better representative, so this page should point towards or redirect to it where appropriate.
- Access without organic visibility: the page is useful in the customer journey but does not need to appear independently in search.
- Unclear: the page’s purpose, audience or relationship with other URLs needs resolving first.
This avoids a common mistake: treating indexing as the goal in itself. Inclusion is not ranking, impressions, traffic or revenue. Establish whether inclusion is commercially justified first.
A practical investigation sequence
1. Confirm the exact URL and its current state
Start with the URL that Search Console actually reports. Check the protocol, hostname, trailing slash, capitalisation, parameters and redirects. Many investigations begin with one URL and end up inspecting a slightly different one.
Then check the current state in URL Inspection and compare it with the live page. Record:
- the reported indexing status and date, where available;
- the last crawl information;
- the HTTP response and redirect chain;
- the user-declared canonical;
- the Google-selected canonical, if reported;
- any indexing or enhancement warnings;
- whether the page has changed since Google last crawled it.
A technical problem becomes more plausible if the URL now redirects unexpectedly, returns an error, has changed content, or carries an indexing directive that conflicts with the intended outcome.
There are other possibilities. Search Console may be showing an older crawl while the page has recently changed, or the live test and coverage report may describe different processing moments.
Do not make a site-wide change based on the label alone. Establish which URL and which version of the page you are diagnosing. If the page has just been released or materially changed, record a review point rather than repeatedly requesting indexing.
2. Check whether the page is technically eligible
Next, check the signals that determine whether Google is allowed and able to treat the page as an indexable document. Look at the robots meta tag, the X-Robots-Tag HTTP header, response status, redirects, access restrictions and the rendered page. Google’s crawling troubleshooting guidance helps separate fetch and page-state problems from assumptions about indexing.
A page can return a successful response and still fail to provide a usable document. JavaScript may not load the main content. A template may render an empty state. Critical resources may fail. The page may also contain a noindex directive added by a CMS rule or release.
Look for a noindex directive when the page should be indexable, missing intended content in the rendered HTML, an unexpected redirect, or a shared template that produces an invalid or incomplete page. Those findings support a technical diagnosis.
If the checks are clean, Google may have selected another representative URL, assessed the page as too similar to an alternative, or simply not completed processing. Fix a confirmed technical constraint, but do not add generic copy merely because the page is not indexed. Technical eligibility is necessary; it does not guarantee inclusion.
3. Compare the declared canonical with likely alternatives
A canonical is the URL a site declares as the preferred representative when several URLs contain the same or substantially similar content. Check the page’s canonical tag, redirects, internal links and sitemap entry, then compare them with Google’s selected canonical where that information is available.
Look for realistic alternatives, including:
- parameter or filtered versions of the same page;
- print, tracking or campaign URLs;
- old and new versions after a migration;
- regional or language variants with overlapping content;
- product, category or service pages that repeat the same core information;
- URLs produced by different CMS routes for the same entity.
Google’s documentation on consolidating duplicate URLs makes an important point: canonical tags, redirects, internal links and sitemap inclusion are signals, not guarantees. Google can select a different canonical when its assessment of the pages points elsewhere.
Consolidation becomes the leading explanation when the inspected URL and another URL have substantially the same purpose and content, Google reports the other URL as canonical, redirects or internal links consistently favour the alternative, or the site repeatedly presents both URLs as the same page.
Similar-looking templates do not necessarily mean duplicate intent. The pages may serve genuinely different needs, or the canonical data may be stale while Google processes the latest version.
Consolidate only when the relationship is real. Depending on the case, that may mean a redirect, a consistent canonical signal, stronger internal linking to the preferred URL, or a change to the information architecture. A self-referencing canonical is not proof that Google must select that URL.
When Google-selected canonical data is unavailable or unclear, compare the target URL with the most likely alternatives rather than treating the absence of a tool verdict as proof of anything.
4. Test the page’s distinct purpose and value
Now move from technical eligibility to editorial and commercial judgement. Ask what a visitor gets from this URL that they cannot get from the likely canonical or a competing page.
For example, a travel site might have separate pages for two genuinely different railway stations, even if their templates are similar. It may not need ten almost identical pages for nearby suburbs if each contains the same transport information and offers no local distinction.
Compare the pages as a user would:
- Do they answer different questions?
- Do they represent different products, services, locations or entities?
- Does the page contain first-party information or functionality specific to its subject?
- Would a customer be disappointed if Google sent them to the alternative instead?
- Does the page have a clear role in the site’s architecture and commercial journey?
Insufficient distinct value is more plausible when a page substantially repeats another URL, has only a swapped place name or product attribute, exists because the CMS generated it, or has no clear user need separate from an existing result.
Google’s helpful-content guidance is qualitative. It does not establish a universal word-count, similarity percentage or search-demand threshold for inclusion.
A short page can still serve a distinct purpose. Length is not a reliable substitute for usefulness, and the absence of a published threshold means this remains a judgement based on user need, page purpose and competing alternatives.
Improve the page only when there is a genuine user need to serve. Add specific information, functionality, evidence or merchandising that makes it meaningfully useful. Do not pad it with generic paragraphs to create the appearance of uniqueness. If the page has no defensible separate purpose, consolidation or intentional exclusion may be the better decision.
5. Check for soft-error behaviour
A soft 404 is a page that returns a technically successful response, often HTTP 200, but appears to Google to be missing, empty, unavailable or erroneous. Google’s crawling troubleshooting guidance describes this distinction.
Check both the raw response and the rendered experience. Look for:
- “product unavailable”, “profile not found” or similar messages;
- empty templates caused by missing database data;
- generic error content returned with a 200 response;
- failed JavaScript that removes the main content;
- pages that look complete to a crawler but broken to a user, or vice versa;
- an expired item that should return a real 404 or redirect to a genuinely relevant alternative.
Soft-error treatment becomes more likely when the page’s primary purpose is unavailable, the rendered page is effectively an error state, or many URLs in the same data set return the same empty template.
There are exceptions. A page may be intentionally sparse but valid, such as a simple event date, stock status or reference page. A 200 response does not prove that a page is healthy, but a short page is not automatically a soft 404 either.
Repair the data or template when the page should exist. Return an appropriate status, redirect or replacement when it should not. For a fuller diagnostic framework, see our guide to soft 404s at scale.
6. Review internal links and sitemap inclusion as supporting evidence
Internal links help Google discover pages and understand the site’s structure. Sitemaps provide a list of URLs that a site considers important. Both are useful signals, but neither guarantees crawling, indexing or ranking. Google explains the role of crawlable links and sitemaps separately.
Check whether the URL:
- has relevant links from pages that are themselves accessible and indexed;
- is buried several clicks deep or only linked from weak navigation paths;
- appears in the correct sitemap;
- has a sitemap
lastmoddate that reflects a meaningful change; - is treated consistently with similar pages that are indexed.
An architecture problem is more plausible if the page is effectively orphaned, the template fails to link to an entire product or location family, or healthy control pages receive materially stronger internal support.
Because the URL has already been crawled, weak discovery is unlikely to be the immediate cause. The real issue may still be duplication, rendering, page purpose or normal processing.
Improve internal linking or sitemap accuracy when they do not reflect the intended architecture. Do not add a URL to a sitemap to force inclusion, and do not resubmit the same sitemap repeatedly without new evidence.
7. Allow for processing time, but make waiting a decision
New or recently changed pages can take days or weeks to move through crawling and indexing. Google states that processing can take time, and that requesting a recrawl does not guarantee inclusion or make repeated requests faster.
Waiting is reasonable when the page is recent, technically accessible, internally linked, clearly distinct, correctly represented in the sitemap and not competing with an obvious alternative. Comparable pages published around the same time may also be waiting.
It is a weaker explanation when the page has been live for a substantial period, similar pages are indexed, or the same exclusion affects a whole template. In those cases, “wait” can become a way of avoiding diagnosis.
Set a review date based on the page type and the evidence available. There is no universal indexing service-level agreement. If the page remains excluded at the review point, compare it again with healthy controls and revisit the stronger hypotheses.
Investigate the page family, not just the URL
One URL can be an isolated oddity. A pattern across hundreds of URLs is usually more useful evidence.
Group affected pages by product set, location set, template, CMS type, language, data feed or release date. Then select a small comparison sample:
- several affected URLs;
- recently published URLs with the same template;
- known healthy URLs from the same family;
- one or two pages that serve a similar purpose but are indexed.
Compare response codes, rendered HTML, content availability, canonicals, redirects, internal links, sitemap inclusion and the underlying business data. This can reveal a template-level noindex rule, a broken feed, repeated canonicalisation, missing content or a release problem that individual URL checks would obscure.
A repeated exclusion pattern suggests investigating the shared system rather than editing URLs one by one. That is an inference from the pattern, not proof of a technical defect: an entire family may also be intentionally consolidated or may simply not warrant separate search results.
For larger sites, the evidence may need to include a crawl, rendered-HTML comparisons, Search Console exports and, where the implementation warrants it, server-log analysis. These sources can help establish what changed and which URL families are affected, but they cannot reproduce every internal indexing decision.
What should happen next?
Use the investigation to choose one of five outcomes:
- Fix: resolve a confirmed directive, redirect, rendering, response or data problem, then validate the affected template and a sample of URLs.
- Consolidate: make the preferred representative clear when two URLs serve the same purpose, using the appropriate canonical, redirect and internal-link signals.
- Improve: strengthen a page only when it has a real, separate user purpose that is currently poorly served.
- Intentionally exclude: keep the page available where useful, but accept that it does not need its own organic result.
- Wait and monitor: use this path when the page is recent and the stronger technical, canonical, rendering and purpose checks are clean.
Requesting indexing, adding more words or resubmitting a sitemap may be part of a sensible workflow, but none guarantees inclusion. More importantly, none replaces the decision about whether the URL should be indexed at all.
The useful distinction is diagnosis, not status chasing
“Crawled – currently not indexed” is a starting point. It tells you Google has visited the URL and that the URL is not currently represented in the relevant index state. It does not tell you whether the cause is duplication, canonical selection, a soft error, weak distinct value or processing time.
The practical method is to move through the evidence in order: confirm the exact URL, check its current technical state, compare it with likely alternatives, test its distinct purpose, inspect soft-error behaviour, review supporting architecture signals and then decide whether waiting is reasonable.
On a small site, a marketer can often investigate one straightforward URL in-house. Ambiguous patterns across many templates are different. They may require a crawl, rendered-HTML analysis, sitemap and Search Console data, and sometimes server logs, alongside someone who can connect the technical evidence to commercial page purpose.
That is where specialist SEO support becomes useful: not to promise that every URL will be indexed, but to diagnose the pattern, prioritise the pages that genuinely matter and help implement the safest response.
Share this article