Your sitemap is how engines discover every page. Paste a domain — we find the sitemap, count URLs, flag duplicates, missing lastmod and the 50k limit, and spot-check a few URLs for 200/404.
A sitemap.xml checker validates the XML file that tells search engines and AI crawlers which URLs on your site exist and when they last changed. This tool counts the listed URLs, flags duplicates and missing lastmod dates, tests the file against the 50,000-URL / 50 MB protocol limit, and spot-checks whether listed URLs actually return 200 instead of 404. A clean, accurate sitemap is how crawlers discover and prioritize your pages — a broken or bloated one wastes crawl budget and leaves content undiscovered, which also keeps it out of AI answers.
It's usually at yourdomain.com/sitemap.xml or referenced in robots.txt on a Sitemap: line. Large sites use a sitemap index file that points to several child sitemaps.
A single sitemap file is capped at 50,000 URLs and 50 MB uncompressed. Beyond that, split your URLs across multiple sitemaps and list them in a sitemap index file.
Yes, when it's honest — Google uses an accurate <lastmod> to decide re-crawl priority. But consistently wrong dates, like today's date on every URL, get ignored and can erode trust in the file.
Listing URLs that return 404 or redirect signals a stale, low-quality sitemap, wasting crawl budget and slowing discovery of your real pages. Keep the sitemap to live, canonical, 200-status URLs only.
Free tools show you WHAT to fix. The full audit + the autonomous team of 20 AI agents fix it end-to-end — strategy, content, GEO, internal links, publishing and day-2 upkeep.
All free tools →cookies
The Matrix already knows everything about you — cookies are small change by comparison. The choice, as always, is yours: