NEW CASE
Antana

We create digital solutions that work for businesses


Give us a call +38 (066) 35-14-529

Let's take the first step towards your website — write to us

Close
Tools

Sitemap Validator

XML validity, duplicates, redirects and whether the URLs respond.

What the validator checks

A broken sitemap is worse than none: Google stops trusting it and crawls the site as if the file did not exist.

XML correctness An unclosed tag, a stray character at the start, an unescaped ampersand in a URL — any of these makes the file unreadable.
Protocol compliance The right root element and the sitemaps.org namespace. Without it some systems ignore the file.
Duplicates and foreign URLs The same page twice, a link to another domain, relative instead of absolute URLs — all are format errors.
Whether pages respond We open a sample of URLs and read the status. A 404 or a redirect has no place in a sitemap.
Conflict with noindex A page listed in the sitemap yet closed to indexing. To Google that signals the settings contradict themselves.
Canonical points elsewhere If a page names another version as primary, listing it in the sitemap makes no sense.

URLs are checked on a sample basis — the first fifteen. Crawling thousands of pages in one request is not possible, so the rest are counted and analysed by format only.

Format limits

These are requirements, not suggestions: exceeding them invalidates the file.

50,000 URLs The maximum per file. Beyond that, split into several and join them with an index.
50 MB The uncompressed size limit. Gzip compression is allowed and encouraged.
One domain A sitemap covers only the domain it sits on. URLs from other sites are ignored.
Date format The lastmod field follows ISO: 2026-01-15, optionally with a time. A familiar 15/01/2026 is simply ignored.
priority Google has ignored this field for years. Setting priorities is not worth the effort.

The most frequent mistakes

URLs blocked in robots.txt The sitemap says “crawl this”, robots.txt says “stay out”. Google sees the contradiction and trusts both files less.
A stale sitemap after a migration Half the URLs redirect to new ones. It technically works while wasting crawl budget.
Duplicates from a trailing slash /page and /page/ are two different URLs to a sitemap. Include the version the site serves with a 200.
An unescaped ampersand In URLs with parameters the & must be written as &. Otherwise the XML breaks at the first such URL.
A blank line before the XML A space or newline before the first tag breaks parsing. A common side effect of generating the file with a script.
The sitemap is never updated Generated once a year ago. New pages never appear, deleted ones stay — such a file is of little use.

What to do with the result

1 Format errors first If the XML is invalid nothing else matters: the file is unreadable and not a single URL counts.
2 Remove URLs that do not belong Redirects, 404s, noindex pages and anything blocked in robots.txt do not belong. What remains is what should be in search.
3 Automate the updates The sitemap should regenerate whenever pages change. Manual updates are forgotten within a month.
4 Submit it in Search Console It shows how many URLs Google read and how many it indexed. The gap between those numbers is the most useful signal you get.
Telegram Viber Call us