Soft 404s at scale: a page-state framework for diagnosis and remediation

A practical framework for finding pages that return 200 but are empty, unavailable or incorrectly rendered. Learn how to combine crawl, rendering, search, site-graph and business-state evidence before choosing a remediation path.

A page can return a technically successful 200 response and still fail its main job. It may contain no meaningful content, show an expired product, return an empty category, fail after JavaScript execution or display a generic success page for a resource that no longer exists.

That creates a diagnostic problem at scale. HTTP status tells you how the server handled a request; it does not establish that the requested resource is useful, available or correctly rendered. Google documents that pages returning 200 can still be recognised as soft 404s when their content indicates that the resource is missing, empty or unusable. Its examples include empty pages, empty internal-search results, broken database connections and prominent error messages (Google’s crawling error guidance).

This article sets out a practical method for investigating that gap. The aim is not to label every thin page a soft 404. It is to classify each URL using several evidence sources, identify whether the cause sits in the CMS, feed, routing layer, rendering stack or SEO controls, then choose an outcome: recover, retain, redirect, remove, apply noindex or investigate further.

Start with page state, not HTTP status

The useful question is not “Does this URL return 200?” It is “What state is this URL in, and does that state match its intended purpose?”

HTTP status remains important. A permanently unavailable resource with no meaningful replacement should generally return 404 or 410, rather than presenting a generic successful page (Google’s soft-404 documentation; HTTP Semantics). Status is one observation within a larger system, not a complete diagnosis.

For a large website, build a Page-State Evidence Matrix. This is a Plus IQ diagnostic model, not a documented Google classifier. For every sampled URL, record six dimensions:

  • HTTP state: status code, redirects, canonical response and response timing.
  • Delivered content: raw HTML, title, headings, body text, structured data and response size.
  • Rendered state: post-JavaScript DOM, screenshot, resource failures, console errors and API responses.
  • Site-graph signals: internal links, navigation placement, orphan status and XML sitemap membership.
  • Search signals: Search Console indexing or soft-404 observations, URL Inspection findings, impressions, clicks and crawl data.
  • Business state: whether the entity exists, is temporarily unavailable, has expired, is suppressed, has returned or should never have been published.

The output should be a decision state rather than a binary label. Useful states include recover, retain as temporarily unavailable, redirect, remove with 404 or 410, retain with noindex, investigate infrastructure and retain as a legitimate sparse page.

Build the investigation around patterns

Start with a representative sample rather than deciding on every URL independently. Segment the population by template, URL pattern, business state, language or market, sitemap inclusion, status code and last meaningful update.

Include healthy control URLs alongside suspected failures. For example, sample an available product, a temporarily unavailable product and a permanently discontinued product from the same template. For a category template, compare populated categories with deliberately empty categories and categories that became empty because a feed or filter failed.

This comparison helps separate a page-level symptom from a template-level cause. If hundreds of affected URLs share the same response body, title pattern, missing component or API error, the remedy is unlikely to be a manual URL exercise. The cause may sit in the CMS, product feed, category logic, rendering layer or routing rules.

Response similarity and random-invalid-URL comparisons can help prioritise soft-error investigations, particularly where invalid URLs receive a standard application response. Academic work has explored response similarity and redirection classification for this purpose, but these methods depend on assumptions about the host. They can produce false positives or miss application-level failures (research on soft errors and web decay; research on soft-error detection by redirection classification). Use them to prioritise investigation, not as the final decision rule.

Classify the main failure patterns

1. Empty or thin templates

Some URLs return the expected template but contain little or no entity-specific content. Common causes include a failed CMS publication, a missing database record, an incomplete migration or a template that renders navigation and boilerplate while omitting the main content container.

Inspect the raw response for the expected title, heading, body fields, structured data and primary content container. Compare it with a healthy URL from the same template. A small response, low word count or high template similarity is a useful triage signal, but it does not prove that a page is a soft 404. A short glossary definition, event page, location page or historical reference can be legitimate and useful. Google documents content-deficient pages as possible soft 404s, but does not publish universal word-count, byte-size or similarity thresholds (Google’s soft-404 guidance).

Check whether the page has a genuine reason to exist. If the underlying record exists and the content has been accidentally suppressed, recover it. If the URL was created without a valid entity, remove it. If the page is intentionally sparse but serves a clear user need, retain it and monitor it rather than applying an arbitrary content threshold.

2. Expired or unavailable entities

An unavailable entity is not automatically a missing entity. A product may be temporarily out of stock, a conference may have ended, a travel property may be closed for a season or a software plan may have been retired while its documentation remains useful.

Combine the page response with business-state data: inventory or availability feeds, publication dates, expiry flags, replacement IDs, booking windows and editorial ownership. Google’s ecommerce guidance supports retaining useful temporarily unavailable product pages when the page clearly communicates its state. It also recognises that an unavailable page is not necessarily appropriate for continued indexing in every situation (guidance on pausing an online business; Google’s ecommerce guidance).

For non-ecommerce entities, apply that principle cautiously. A useful historical article or completed event may deserve retention for reference. A permanently expired listing with no remaining information may not. Business state and user purpose matter more than the age of the URL.

3. Zero-result categories

A category or listing page with zero results can represent several different states:

  • There are genuinely no current entities, but the category is expected to return in future.
  • A seasonal or regional inventory state has temporarily emptied the page.
  • A feed, taxonomy or filter has failed.
  • The category was created accidentally and has no meaningful purpose.
  • The page is useful as an editorial or navigational destination despite having no current listings.

Do not treat every empty category as an automatic removal candidate. Compare the URL with business calendars, feed data, historical availability and equivalent categories in other markets. Google’s ecommerce documentation describes state-based handling for empty or unavailable category experiences, but the correct outcome still depends on future purpose, user value and site behaviour (Google’s guidance on ecommerce URL structures).

If a category should contain results but does not, fix the data or category logic first. If it is a valid but temporarily empty destination, retain it with a clear explanation where that serves users. If it has no future purpose, remove it. If it remains useful to users but is not intended for organic search, noindex may be appropriate, provided the page remains crawlable so search engines can see the directive (Google’s robots meta tag documentation).

4. JavaScript failure states

JavaScript can create a soft-404-like state in either direction. The raw response may look empty even though the page should populate after execution. Alternatively, the raw HTML may contain a valid shell, but client-side code can replace it with “no results”, an error message or an empty component after an API request fails.

Google documents a crawling, rendering and indexing process for JavaScript-dependent pages. For pages whose meaningful content is generated or changed in the browser, raw-response inspection alone is therefore insufficient (JavaScript SEO basics).

For a representative sample, compare:

  • the initial HTML response with the post-render DOM;
  • the rendered screenshot with the expected user experience;
  • network requests, API responses and HTTP errors;
  • console errors, blocked resources and timeout behaviour;
  • output across relevant user agents, devices, locales, consent states and logged-in or logged-out contexts.

Google’s troubleshooting guidance supports checking rendered output and JavaScript-related failures, but a single test context cannot prove that every crawler or user receives the same result (fixing JavaScript search issues). A repeated failure across a template is an implementation problem, not simply an SEO classification problem.

5. Misleading success responses

The clearest example is a missing URL that returns the site’s generic homepage, search page or “request received” template with a 200 response. Another is a blanket redirect from discontinued pages to an unrelated category.

A redirect is appropriate when the destination is genuinely equivalent. It is not a universal substitute for a missing resource. Google warns against redirecting missing URLs to irrelevant destinations or the homepage because the result can be misleading and may be treated like a soft 404 (Google’s redirect guidance).

Test invalid URL controls deliberately. Request a set of clearly non-existent paths and compare their status, title, body, canonical, internal links and rendered output with the suspected URLs. If both receive the same “successful” page, the routing layer is likely masking the distinction between valid and invalid resources.

Connect the evidence layers

Each evidence source answers a different question:

  • HTTP and raw HTML: What did the server deliver, and did it distinguish the resource correctly?
  • Rendered DOM: What does the page become after scripts, API calls and client-side state have run?
  • Internal links: Does the site still present this URL as part of its navigable information architecture?
  • XML sitemaps: Is the publisher still communicating that the URL is intended for discovery?
  • Server logs: Are crawlers and users requesting the URL, and what responses are they receiving?
  • Search Console: Has Google observed a soft 404, excluded the page or shown search demand?
  • Business data: Does the underlying entity exist, and what is its current or future state?

Sitemap membership is useful as an intent signal, but it does not guarantee crawling, indexing or canonical selection (sitemap overview; building and managing sitemaps). Internal links and sitemap inclusion can also be stale, so disagreement between them is evidence to investigate rather than a verdict.

Search Console adds Google’s processed view, including indexing and soft-404 observations, but it is delayed and not a complete real-time inventory. Combine it with a live crawl, logs and current business data rather than treating one report as definitive (Search Console indexing reports).

Map evidence to the right outcome

Once a URL has been classified, map it to an outcome that reflects its actual state:

  • Recover: repair the CMS record, feed, category logic, API or rendering path when the intended resource still exists.
  • Retain as temporarily unavailable: keep a useful page when the entity may return or its information remains valuable. State the availability clearly and avoid presenting an unavailable action as complete.
  • Redirect: use a relevant, genuinely equivalent destination when the original resource has been replaced. Document the mapping rather than applying a blanket homepage rule.
  • Remove with 404 or 410: use this when the resource is permanently gone and there is no meaningful replacement. The evidence does not establish a universal SEO advantage for choosing 410 over 404; match the implementation to the actual resource state.
  • Retain with noindex: use this for an accessible page with user or operational value that is deliberately excluded from organic search. Do not use it to hide a routing or content failure, and ensure crawlers can access the page to see the directive (noindex documentation).
  • Retain as legitimate sparse content: keep a short page when it has a distinct purpose, clear information and evidence of user or business value.
  • Investigate infrastructure: use this when the state varies by user agent, timing, geography or rendering context, or when the evidence points to an intermittent platform failure.

The difficult part is rarely finding another unusual URL. It is deciding whether the page is missing, temporarily unavailable, incorrectly rendered or intentionally sparse, then assigning ownership for the fix.

Synthetic example: a travel publisher’s expired destination pages

Imagine a travel publisher with thousands of pages for hotels and destinations. A crawl finds 18,000 pages returning 200 with little visible text. A simple word-count rule would classify them all as failures.

The evidence matrix produces a more useful split. Some pages are valid destination guides with short but complete factual content. Some hotel pages have an active record but a failed availability API, leaving the rendered booking module empty. Others refer to properties that have permanently left the publisher’s inventory. A fourth group remains in the XML sitemap and is linked from old destination pages, although the underlying records were deleted.

The remediation differs by group. The API failure belongs to the rendering or integration team. Temporarily unavailable properties may remain live with a clear state and useful information. Permanently deleted properties need a relevant replacement only where one genuinely exists; otherwise they should return an appropriate not-found response. The sparse destination guides should be retained if they answer a distinct travel question.

The useful finding is not that 18,000 pages are “thin”. It is that one apparent symptom contains several business and technical states.

Implement and QA the changes

Remediation should be released as a controlled change, not as a bulk status-code edit.

  1. Define the populations: group URLs by template, business state and suspected failure pattern. Preserve healthy and invalid control samples.
  2. Assign ownership: route CMS, feed, taxonomy, rendering, routing and SEO-control issues to the team that can change the cause.
  3. Implement the intended state: change content, availability messaging, API handling, redirects, sitemap inclusion, internal links or response codes as appropriate.
  4. Test representative URLs: check each template and outcome across raw HTML, rendered DOM, status, canonical, robots directives, structured data and user-facing messaging.
  5. Check deployment boundaries: verify caches, CDNs, edge rules, localisation, authentication, consent states and mobile or desktop variants. These contexts can change the delivered or rendered result.
  6. Recrawl: validate a sample immediately, then recrawl the affected population after the change has propagated.
  7. Monitor recurrence: track new empty responses, invalid URLs receiving successful templates, rendering failures, unexpected sitemap membership and changes in business-state alignment.

A useful internal monitoring proposal is the proportion of sampled URLs where HTTP, rendered, site-graph and business states disagree. This is a suggested diagnostic metric, not an established industry standard. Define the states, establish a baseline and validate that changes in the measure correspond to real user or search problems before using it as a performance target.

For related work on how navigation changes affect crawl paths, see the analysis of JavaScript navigation and crawl-graph drift and faceted navigation and internal-link inflation.

Keep the limitations visible

Soft-404 classification is partly opaque. Google does not publish a complete reproducible classifier or universal thresholds for emptiness, thinness or response similarity. Search treatment can change, and Search Console data represents a processed view rather than every current request.

HTTP status alone is insufficient, but the opposite error is also possible: a sparse or unavailable page is not automatically wrong. User value, historical importance, future availability and business purpose can justify retaining a page. Rendering evidence can vary by context, and business-state data can itself be stale. The Page-State Evidence Matrix therefore supports prioritisation and judgement; it does not remove the need for representative testing.

Conclusion

The central distinction is between a URL that is technically reachable and a resource that is genuinely available and useful.

At scale, that distinction cannot be established from 200 responses, word counts or Search Console labels alone. Compare transport, content, rendering, site-graph, search and business-state evidence. Then give the URL the state it actually deserves: recovery, retention, redirection, removal, noindex or further investigation.

That approach turns soft-404 work from a hunt for suspicious pages into systems diagnosis. It also makes remediation safer: legitimate sparse pages are less likely to be removed, while CMS, feed, routing and rendering failures are more likely to reach the teams responsible for fixing them.

Share this article

Found this useful? Pass it on.

Share on LinkedIn · Share on X