# OpenScraper Sitemap Scraper API

*HTML page: https://openscraper.ai/sitemap-scraper-api · Agent index: https://openscraper.ai/llms.txt*

> Crawl any website's sitemap(s) and get every URL back as clean, structured JSON — with lastmod, priority, and change frequency included. For $0.01 per 1,000 URLs plus $0.002 per sitemap crawled.

## What it does

A `sitemap.xml` lists every URL a website wants search engines to find — the fastest way to discover a site's full page inventory without crawling links one by one. Our Sitemap Scraper API finds every sitemap a site publishes (via `robots.txt` and sitemap indexes) and returns every URL it contains, complete with `lastmod`, `priority` and `changefreq`.

Give it a domain and the crawler does the rest: it locates the root sitemap (or sitemap index), follows every nested sitemap it references, and streams back every URL as clean, structured JSON — ready for SEO audits, content migrations, competitive research, or feeding a scraper of your own.

Already know which sitemap you want? Pass its URL (`https://example.com/sitemap-fr.xml`) instead of the domain and only that document is crawled — no `robots.txt`, none of the site's other sitemaps.

- From $0.01 per 1,000 URLs + $0.002 per sitemap crawled
- Structured JSON output
- Nested sitemap indexes handled
- Residential IP bypass available
- Works on any website, worldwide

Sites that block datacenter IPs can be crawled through our residential IP pool on request, and there is no cap on URLs per run.

## Output fields

`url`, `sitemap_url`, `site`, `lastmod`, `priority`, `changefreq`, `position`, `entry_type`, `search_string`, `scraped_at`.

`lastmod`, `priority`, `changefreq` and `position` are always present, and `null` when the sitemap declares nothing for them.

## Pricing

| Product | Price without subscription | Price with subscription |
| --- | --- | --- |
| URLs | $0.01 per 1,000 URLs | $0.0085 per 1,000 URLs |
| Sitemaps | $0.002 per sitemap crawled | $0.0017 per sitemap crawled |

A run is billed on both: every sitemap document the crawl downloads, and every URL it returns.

Residential IPs (for sites that block datacenter IPs): 4× both rates, opt-in per run.

Plans and credit rules: <https://openscraper.ai/pricing.md>.

## Using it from the API

The crawl returns many rows, so run it asynchronously and poll (or stream the full result set with `GET /runs/:id/export.json`).

```bash
curl -X POST https://api.openscraper.ai/runs \
  -H "Authorization: Bearer sk_live_xxx" \
  -H "Content-Type: application/json" \
  -d '{
    "module": "sitemap_matrix",
    "params": {
      "url": "https://www.mercedes-benz.com/",
      "max_results": 50000,
      "crawl_timeout": 60,
      "concurrency": 15,
      "use_residential": false
    }
  }'
```

| Module | Purpose | Key parameters |
| --- | --- | --- |
| `sitemap_matrix` | Discover and crawl a site's sitemap(s) | `url` (root domain), `max_results` (no ceiling), `crawl_timeout`, `concurrency`, `use_residential` |

Full parameter and response reference: <https://openscraper.ai/docs.md>.
