Which URLs Belong in an XML Sitemap? A Practical Inclusion Guide
A practical guide to deciding which URLs belong in an XML sitemap, why preferred indexable pages usually qualify, and what sitemap inclusion cannot guarantee.
The question is not simply whether a URL exists. It is whether that URL is the version you want search engines to discover, understand and potentially show to searchers.
That distinction matters because websites often contain technically valid URLs that are poor sitemap candidates: tracking-parameter versions, redirects, duplicate pages, noindex campaign pages and URLs used for internal workflows rather than organic search.
This guide sets out a practical inclusion policy. In most cases, a sitemap should contain the preferred URL that returns a usable page, is intended to be indexable and represents a worthwhile search destination. A sitemap can help search engines discover URLs and communicate which versions you consider important. It cannot force crawling, indexing or rankings.
Three questions to ask about every URL
Sitemap decisions become clearer when you separate three questions that are often treated as one:
- Is the URL important to the business? Does it support revenue, leads, customer education, a product launch or another business objective?
- Is the URL eligible for organic search? Is it intended to appear in search, rather than being private, temporary, duplicated, blocked or explicitly marked
noindex? - Is it appropriate for the sitemap? Is it the preferred version of a useful search destination that you want search engines to discover?
These decisions are related, but they are not interchangeable. A campaign landing page may be important to the marketing team but deliberately carry a noindex directive. A customer portal may be essential to the business but have no place in an organic search sitemap. Conversely, a newly published advice page may be suitable for the sitemap before it has attracted much traffic.
This is an applied SEO policy rather than a formal three-part framework set out by Google. It follows from how Google describes sitemaps, indexing controls and preferred URL signals. Business importance or technical reachability alone does not make a URL suitable for the sitemap.
The practical rule: include the preferred, useful, indexable URL
For most sites, a URL is a good sitemap candidate when it passes four checks:
- It is the preferred version. It is not a duplicate, tracking variant or alternate URL that should consolidate to another page.
- It returns a usable response. In normal circumstances, the URL resolves to the page rather than redirecting, failing or returning an error.
- It is intended to be indexable. The page is not deliberately excluded with a
noindexdirective or another site policy. - It represents a worthwhile search destination. It has a distinct reason to exist for a searcher and fits the site's organic search strategy.
A 200 response is only one part of the decision. It tells you that the server returned a successful response; it does not tell you whether the page is canonical, unique, indexable or useful for search. The HTTP specification defines response status, while Google's sitemap guidance and canonicalisation guidance provide context for deciding whether the URL is a sensible preferred destination.
A sitemap is therefore not a list of every URL the CMS can produce. It is closer to a list of the search destinations you would be comfortable recommending to a crawler.
What usually belongs in an XML sitemap?
Preferred product, category and service pages
Core commercial pages will usually belong in the sitemap when they are unique, indexable and intended to attract organic visibility. That might include a product URL, a category page, a location page or a clearly differentiated service page.
For example, if /running-shoes/ is the preferred category URL and returns the page you want customers to find, it is a sensible sitemap entry. The same applies to a product such as /products/trail-runner-x/, provided that the product page is not a duplicate of another URL and your search policy allows it to be indexed.
Google recommends sitemaps as a way to help search engines discover URLs, particularly on large, complex, new or poorly linked sites. A clean set of preferred commercial URLs can therefore be useful even when those pages are also linked through navigation and internal content. Google's sitemap overview explains these discovery use cases.
Useful content pages with a genuine search purpose
Editorial pages, guides, research pages and help content can belong in the sitemap when they offer a distinct answer or resource and are intended to appear in search.
Suppose a company publishes a useful guide to choosing a home energy tariff. The page has only a few internal links because it was published yesterday, but it is complete, indexable and part of the organic search plan. It can be an appropriate sitemap URL while the marketing team improves its internal linking separately.
That distinction matters. Sitemap inclusion can support discovery, but it does not repair weak site architecture or guarantee that the page will be crawled or indexed. The sitemap is a helpful signpost, not a replacement for a clear route through the website.
Newly published pages ready for search
A useful page does not need to wait until it has earned traffic or accumulated a certain number of links before it enters the sitemap. If it is the preferred URL, returns a usable page, is intended for search and has a clear reason to exist, inclusion is reasonable.
Its absence from the sitemap does not automatically create a defect. Search engines may discover and index an important page through internal links or other references. The omission becomes more concerning when the page is important, search-eligible and difficult to discover, especially on a large, new or frequently changing website.
What usually does not belong?
Redirecting URLs
A URL that redirects should normally be excluded in favour of the final destination. If /old-running-shoes/ redirects to /running-shoes/, the latter is usually the sitemap URL.
A redirect is useful for preserving old links or supporting a migration, but it is not normally the page you want to present as a current search destination. The HTTP specification's redirect definitions and Google's guidance on moving URLs support this distinction.
Duplicate and tracking-parameter URLs
Imagine a newsletter link that produces:
https://example.com/products/trail-runner-x/?utm_source=newsletter
If that URL displays the same product as the clean version, it should usually stay out of the sitemap. List the preferred product URL instead:
https://example.com/products/trail-runner-x/
This is an applied inclusion rule based on Google's advice to consolidate duplicate URL versions and use consistent preferred signals. The source does not define a universal sitemap policy for every parameter, but it supports a practical principle: do not use the sitemap to promote a URL that your site treats as an alternate representation.
Parameters are not automatically wrong. A parameterised URL may represent genuinely distinct, stable content in some systems. The question is whether it has a deliberate search purpose and is the preferred URL for that content, rather than whether a question mark appears in the address.
Noindex pages
A page with noindex is normally a poor sitemap candidate because the directive says that the page is not intended to appear in search.
Consider a campaign page created for paid traffic. It may be commercially important, carefully designed and live on the main domain, but if the organic search policy is to keep it out of search, adding it to the sitemap sends mixed signals. The coherent operational rule is to keep it out and ensure that its intended treatment is clear.
Google's indexing guidance explains how noindex controls whether a page can appear in search. Sitemap inclusion remains a hint; it does not override that directive or turn a deliberately excluded page into an organic search destination.
Robots.txt-blocked URLs
A URL blocked in robots.txt is also normally a poor sitemap candidate. If a crawler cannot retrieve the page, it may not be able to inspect its content, canonical signal or indexing directives.
There is an important caveat: robots.txt controls crawler access, not guaranteed index exclusion. A blocked URL can sometimes appear in search without its contents being crawled. Robots rules should therefore not be treated as a substitute for an indexing policy. The Robots Exclusion Protocol describes crawler access rules, while Google's indexing guidance explains why blocking access can prevent a crawler from seeing other signals.
Errors, empty states and unwanted search destinations
URLs returning errors, empty results or temporary holding pages do not normally belong in the sitemap. Neither do pages that are technically accessible but have no useful organic search purpose.
Faceted navigation is a common example. A retailer may generate thousands of URLs for combinations such as colour, size and delivery preference. Some combinations may be valuable search destinations; many will not be. The sitemap should reflect the deliberate set of preferred, worthwhile pages, not every combination the filtering system can create.
The same judgement applies to out-of-stock products, short-lived seasonal pages and near-duplicate service pages. Their treatment depends on the site's policy and the search value of each URL. Technical validity alone is not enough.
A simple decision model for sitemap membership
For each candidate URL, work through these questions:
- Is this the URL we want to represent the page? If another URL is canonical or preferred, use that one instead.
- Does it resolve to the intended page? Check the HTTP status and final destination. A redirect or error usually means the URL should not be listed.
- Is the page intended to appear in organic search? Check the indexation policy and any
noindexdirective. - Can search engines access the page? Check for a robots.txt rule that prevents retrieval. Do not confuse access control with indexation control.
- Does the page have a distinct search purpose? Ask what a searcher would gain from landing on this URL rather than another page.
- Would we be comfortable treating this as a preferred search destination? If the answer is no, it probably does not belong in the sitemap.
This is intentionally narrower than a full sitemap audit. It is a membership check: a way to compare the sitemap with the signals and policy that define the site's preferred search destinations.
What sitemap inclusion can and cannot achieve
A sitemap can help search engines discover URLs, particularly where a site is large, new, complex or not comprehensively linked. It can also communicate which URL versions the site considers preferred.
Inclusion is only a hint. Google's documentation states that submitting a sitemap does not guarantee crawling, indexing or ranking, and that sitemap inclusion does not itself increase rankings.
It is also a relatively weak canonicalisation signal. Redirects and rel="canonical" are stronger signals than sitemap inclusion, although all canonical signals are still interpreted alongside other evidence. Google's duplicate URL guidance explains why consistent signals matter.
A clean sitemap cannot rescue contradictory implementation. If the sitemap lists one URL, internal links point to another, redirects lead to a third and canonicals identify a fourth, the search engine still has to resolve the conflict. The sitemap is helpful when it agrees with the rest of the site.
Does every valid URL need to be in the sitemap?
No. A small site with roughly 500 or fewer search-relevant pages may not need a sitemap if its important pages are comprehensively linked. Google makes this point in its sitemap overview, while also noting that sitemaps can be useful in other situations.
That does not mean an omitted URL is automatically safe, or that a sitemap is only for very large websites. It means sitemap membership should be proportionate to the site's size, structure and change rate.
As a website grows, a clean set of preferred URLs becomes more valuable. It gives search engines a clearer discovery aid and gives the organisation a practical reference for its indexation policy. On a small, well-linked site, maintaining the same rule may be simple. On a retailer with thousands of products and filters, the difference between a preferred URL and a merely valid URL becomes much more consequential.
How to check sitemap membership in practice
You do not need to begin with a large technical audit. Start with a sample of sitemap URLs and compare them with the signals that should govern inclusion:
- HTTP status: does the URL return the intended page, or does it redirect or fail?
- Final destination: if it redirects, does the sitemap contain the destination instead?
- Canonical signal: does the page identify itself as the preferred URL, or does it point elsewhere?
- Indexation directive: is there a
noindexinstruction that conflicts with inclusion? - Robots access: can a crawler retrieve the page and inspect its signals?
- Search purpose: is the page a distinct, worthwhile destination rather than a duplicate, internal tool or temporary variant?
You can then reverse the check. Look for important indexable pages that are absent from the sitemap, particularly pages that are newly published, weakly linked or generated through templates. Their absence is not proof of a problem, but it is a useful prompt to check whether discovery depends on a route that is reliable enough.
This membership check should sit alongside, rather than replace, broader sitemap, release and indexation work. A sitemap can be technically valid while still disagreeing with canonicals, templates or the site's search-purpose policy.
Make the rule operational
For a straightforward site, the CMS may be able to apply a simple rule: include pages that are published, canonical, indexable and intended as organic search destinations. This is an implementation choice rather than a requirement in Google's documentation, so someone should still verify that the rule matches the site's commercial and editorial decisions.
Large or frequently changing sites usually need more systematic reconciliation. Sitemap output should agree with templates, canonical tags, HTTP responses, robots directives and indexation policy. Otherwise, a template change can quietly place thousands of redirects, duplicate variants or noindex pages into the sitemap.
The right level of automation depends on the site. The aim is not to list every URL that exists. It is to keep the sitemap aligned with the set of preferred pages the business is genuinely prepared to have discovered and considered for organic search.
The practical takeaway
Start with the page you want searchers to find, not with the URLs your system happens to generate.
In most cases, include the preferred URL that returns a usable page, is intended to be indexable and has a worthwhile search purpose. Exclude redirects, duplicate and tracking variants, noindex pages, blocked URLs, errors and unwanted destinations unless a deliberate site policy says otherwise.
Keep the limits in view. A sitemap can help discovery and reinforce preferred URL signals, but it cannot force crawling, indexing or rankings. On a small, well-linked site, not every eligible page needs to be listed. As the site grows or changes faster, maintaining a clean preferred-URL set becomes a useful technical and commercial discipline.
Where the rules become difficult to maintain, Liquid Silver can help diagnose the gap between sitemap output, site architecture and indexation policy, then prioritise the fixes that matter commercially.
Share this article