# OpenScraper API documentation

*HTML page: https://openscraper.ai/docs · Agent index: https://openscraper.ai/llms.txt*

> A single REST endpoint runs every scraper. Authenticate with an API key, POST the module and its parameters, and get structured data back — synchronously or by polling.

## Introduction

Every scraper on OpenScraper is driven through one HTTP API. The base URL for all requests is:

```text
https://api.openscraper.ai
```

You pick a module (for example the Web Unlocker or the Google Maps scraper), pass its parameters, and the API returns the scraped rows. All requests and responses are JSON, and every request must be authenticated.

## Authentication

Authenticate by sending your API key as a Bearer token in the `Authorization` header:

```http
Authorization: Bearer sk_live_xxxxxxxxxxxxxxxxxxxxxxxx
```

Create and manage keys from your [Dashboard](https://openscraper.ai/dashboard). A key is shown in full **only once**, at creation — store it somewhere safe. Keys start with `sk_live_`.

Keep your API key secret. It bills your account — never expose it in client-side code, a public repo, or a browser request.

## Quickstart

Unlock a page and get its HTML back in a single synchronous call:

```bash
curl -X POST https://api.openscraper.ai/runs \
  -H "Authorization: Bearer sk_live_xxx" \
  -H "Content-Type: application/json" \
  -d '{
    "module": "unlocker_matrix",
    "params": { "url": "https://example.com", "geo": "France" },
    "sync": true
  }'
```

With `sync: true` the API waits for the scrape to finish and returns the result in the same response. For larger jobs, submit asynchronously and poll (see below).

## MCP (Model Context Protocol)

The same API is also exposed as a hosted MCP server, so an AI assistant (Claude on claude.ai, Desktop or Code, Cursor, or any MCP client) can drive it in plain words: probe a site, estimate the price, preview a sample, run the scrape and hand you the results. Same modules, same prices, same credits as the API.

- Server URL: `https://mcp.openscraper.ai/mcp`
- Transport: Streamable HTTP

### Authentication

- **OAuth 2.1** (recommended, for claude.ai, Claude Desktop connectors, Claude Code): the client registers itself, you sign in to OpenScraper and approve. Authorization code + PKCE (S256), dynamic client registration, scope `mcp`. Access tokens last 1 hour and refresh automatically for 90 days. Metadata: `https://mcp.openscraper.ai/.well-known/oauth-authorization-server`.
- **API key**: send `Authorization: Bearer sk_live_...`, exactly as with the REST API. For other clients, scripts and CI.

### Connecting a client

- **claude.ai** (also Claude Desktop and mobile, same account): Settings → Connectors → Add custom connector, name it "OpenScraper", paste the URL, click Connect and sign in.
- **Claude Code**: `claude mcp add --transport http openscraper https://mcp.openscraper.ai/mcp`, then `/mcp` to sign in. With a key instead, add `-H "Authorization: Bearer sk_live_xxx"`.
- **Cursor and clients that take a URL + headers**:

```json
{
  "mcpServers": {
    "openscraper": {
      "url": "https://mcp.openscraper.ai/mcp",
      "headers": { "Authorization": "Bearer sk_live_xxx" }
    }
  }
}
```

- **Stdio-only clients** (e.g. Claude Desktop's config file), through `mcp-remote` (requires Node.js):

```json
{
  "mcpServers": {
    "openscraper": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://mcp.openscraper.ai/mcp",
               "--header", "Authorization:${OPENSCRAPER_AUTH}"],
      "env": { "OPENSCRAPER_AUTH": "Bearer sk_live_xxx" }
    }
  }
}
```

### Tools

| Tool | Arguments | Description |
|---|---|---|
| `list_modules` | — | Dedicated scrapers available to you (Google Maps, Leboncoin, Sitemap…). |
| `probe_site` | url | Reachability, anti-bot detected, engine that gets through, volume estimate, proxy recommendation. Billed as one small probe. |
| `estimate_cost` | module, params | Typical cost and upper bound of a run, in cents. |
| `preview_sample` | module, params, n | Small real sample, synchronous, max 25 rows. Billed per item. |
| `run_scrape` | module, params | Launch a full run (async). Returns a task_id. |
| `get_run` | task_id | Status, live progress and a 5-row preview. |
| `list_runs` | module? | Your recent runs. |
| `cancel_run` | task_id | Stop a run, keeping what was scraped. |
| `export_results` | task_id, format | Download link for all results, xlsx (default) or csv, valid 24 h. |
| `generate_client_code` | module, params, lang | Runnable code to scrape from your own machine. |
| `warm_session` | url, proxy? | Replayable browser session (cookies, user agent, TLS profile). |
| `find_emails` | urls | Contact emails published by websites. |
| `verify_emails` | emails | Deliverability of email addresses. |
| `get_balance` | — | Your credit balance. |

### Typical flow

1. `list_modules`; if no dedicated scraper fits, `probe_site`.
2. `estimate_cost`, then `preview_sample`, confirmed with the user before a big run.
3. Either `run_scrape` + `get_run` (we run it), or `generate_client_code` (you run it).
4. `export_results` for the full data set.

### Large results

Rows never flow through the conversation in bulk: `get_run` shows a 5-row preview, and `export_results` returns a signed link to every row as Excel or CSV (no login, valid 24 hours). An agent with a shell can download the CSV and load it into a database. Sitemap crawls export as JSON through `GET /runs/{id}/export.json`. Over REST: `POST https://api.openscraper.ai/runs/{id}/export-link?format=csv` with your Bearer key.

### Scope and billing

Everything is scoped to the connected account: only your runs, billed to your credits at the usual per-item rates (probes and previews included); an empty balance scrapes nothing. Step-by-step guide: https://openscraper.ai/mcp. Example: https://openscraper.ai/claude-leboncoin.

## Running a scrape

`POST /runs`

Request body fields:

| Parameter | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `module` | string | yes | — | The module to run, e.g. `"unlocker_matrix"` or `"googlemaps_matrix"`. |
| `params` | object | yes | — | Module-specific parameters (see each module's reference below). |
| `sync` | boolean | no | `false` | If true, wait inline for the result. If false, return immediately with a `task_id` to poll. |
| `sync_timeout_seconds` | int | no | 1–600 | Max seconds to wait when `sync` is true. On timeout you get a `task_id` and status `"running"`. |
| `webhook_url` | string | no | — | A URL to be notified at when the run finishes. |

**Synchronous** (`sync: true`) is best for fast, single-item modules like the Web Unlocker. **Asynchronous** (the default) is best for modules that return many rows, like Google Maps.

The run is always billed to the owner of the API key.

## Getting results

### Synchronous response

When the scrape finishes within the timeout, you get HTTP 200 and the full run object:

```json
{
  "id": "6f0b…",
  "module": "unlocker_matrix",
  "status": "done",
  "result": [
    {
      "url": "https://example.com",
      "status_code": 200,
      "content_type": "text/html",
      "result_url": "https://api.openscraper.ai/runs/6f0b…/file"
    }
  ],
  "error": null,
  "progress": { "scraped": 1, "credits_exhausted": false }
}
```

`result` is an array of scraped rows (the fields depend on the module). `status` is one of `pending`, `running`, `done`, `error`, `stopped`.

### Asynchronous and polling

A non-sync submission (or a sync call that times out) returns a `task_id` with status `"pending"` or `"running"`:

```json
{ "task_id": "6f0b…", "status": "pending" }
```

Poll the run until it reaches a terminal status:

```bash
curl https://api.openscraper.ai/runs/6f0b… \
  -H "Authorization: Bearer sk_live_xxx"
```

`GET /runs/{task_id}` returns the same run object shown above. Poll every few seconds until `status` is `done`, `error` or `stopped`. While running, `progress.scraped` tells you how many rows have been collected so far.

## Errors

The API uses standard HTTP status codes:

| Status | Meaning | Description |
| --- | --- | --- |
| `200` | OK | Sync run finished successfully. |
| `202` | Accepted | Run accepted / still running — poll with the `task_id`. |
| `400` | Bad Request | Unknown module or malformed parameters. |
| `401` | Unauthorized | Missing or invalid API key. |
| `403` | Forbidden | Module restricted to admin accounts. |

A run that *starts* but fails mid-scrape lands with `status: "error"` and a short message in the `error` field. Detailed provider-level diagnostics are never returned to clients — they are available in your Dashboard.

## Credits and billing

Billing is pay-as-you-go: you are charged per item actually scraped, debited as the run progresses. There is no upfront reservation.

If your balance cannot cover even one item, the run returns immediately with zero results and `progress.credits_exhausted: true`. If credits run out mid-run, the run stops gracefully and keeps everything scraped so far — it lands as `done`, not `error`.

See per-module rates on the [Pricing page](https://openscraper.ai/pricing).

## Module reference: Web Unlocker

`module: "unlocker_matrix"`

Fetch any URL through our anti-bot stack and get the raw HTML back. Returns exactly one result. Best used synchronously.

| Parameter | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `url` | string | yes | — | The target URL to unlock. |
| `geo` | string | no | `France` | Proxy geolocation used for the request. |
| `headless` | boolean | no | `true` | Render JavaScript in a headless browser before returning the HTML. Disable for a raw, faster HTTP fetch. |

Result fields:

| Field | Type | Description |
| --- | --- | --- |
| `url` | string | The URL that was fetched. |
| `status_code` | int | HTTP status of the upstream fetch. |
| `content_type` | string | MIME type of the returned body. |
| `result_url` | string | A URL to download the raw HTML/body of the page. |

The page body is not inlined in the JSON — fetch it from `result_url` (same Bearer auth).

## Module reference: Google Maps Search

`module: "googlemaps_matrix"`

Export every establishment from a Google Maps search URL — names, addresses, phones, websites, coordinates and more. Returns many rows; run it asynchronously.

| Parameter | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `url` | string | yes | — | A Google Maps search URL. |
| `max_results` | int | no | `200` | Google caps a single search at ~200 results. |
| `collect_contacts` | boolean | no | `true` | Also pull contact details from each place's website. |
| `details` | boolean | no | `false` | Extract extra attributes (`plus_code`, opening hours…). |
| `ratings` | string | no | `Any rating` | Filter places by minimum average rating. |
| `images` | boolean | no | `false` | Extract up to 240 images per listing. |
| `search_country` | string | no | `United States` | Geographic context for the search. |
| `language` | string | yes | `English (US)` | Language of the returned results. |

Each result row includes `title`, `category_name`, `address`, `phone`, `website`, `lat`/`lng`, `place_id` and more.

## Module reference: Google Maps Reviews

`module: "googlemapreviews_matrix"`

Collect all reviews from a Google Maps establishment URL. Run asynchronously and poll.

| Parameter | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `url` | string | yes | — | A Google Maps establishment URL. |
| `sort_by` | string | no | `newest` | Review sort order. `"newest"` collects all reviews. |
| `max_results` | int | no | — | Max reviews to collect per establishment. |
| `hours_back` | int | no | — | Only collect reviews from the last N hours. |
| `language` | string | yes | `English (US)` | Language of the returned reviews. |

Each review includes `user_name`, `score`, `text`, `published_at` and the owner response.

## Module reference: Sitemap Scraper

`module: "sitemap_matrix"`

Walk a site's sitemaps and get every URL they list. Give a root URL and the crawl reads `robots.txt`, then follows every sitemap the site declares (or the usual locations when it declares none), down through sitemap indexes to the pages. Returns many rows; run it asynchronously.

Pass a **sitemap URL** instead (`https://example.com/sitemap-fr.xml`) and only that document is crawled: no `robots.txt`, none of the site's other sitemaps — just the pages it lists, plus the child sitemaps it lists if it is an index.

| Parameter | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `url` | string | yes | — | A site root URL, or a single sitemap URL to crawl only that one. |
| `crawl_timeout` | int | no | `60` | Budget for the whole walk, in seconds. Max 900 — raise it for very large sites. |
| `concurrency` | int | no | `15` | Sitemaps downloaded in parallel. Max 50. |
| `max_results` | int | no | `50000` | Stop after this many URLs. No ceiling — what ends a very large run is `crawl_timeout`, the worker's 300s job timeout, or your credit balance. |
| `use_residential` | boolean | no | `false` | Allow residential IPs when the site turns our datacenter IPs away (403/429). Billed at 4x for the whole run — both the per-sitemap and the per-URL rate. |

Result fields:

| Field | Type | Description |
| --- | --- | --- |
| `url` | string | The page URL found — or, on an error row, the sitemap that failed. |
| `entry_type` | string | `"page"` for a discovered URL, `"error"` for a sitemap that could not be fetched or parsed. |
| `sitemap_url` | string | The sitemap the URL was listed in. |
| `site` | string | Host of the crawled site. |
| `position` | int | 0-based rank of the page inside its own sitemap. |
| `lastmod` | string | `<lastmod>` as the sitemap published it (a W3C datetime). `null` when not declared. |
| `changefreq` | string | `<changefreq>`, lowercased: `always`, `hourly`, `daily`, `weekly`, `monthly`, `yearly`, `never`. `null` when not declared. |
| `priority` | float | `<priority>` as declared, 0.0-1.0. `null` when not declared. |
| `error` | string | Error rows only: why the sitemap failed, with `error_status` carrying the HTTP status. |

`position`, `lastmod`, `changefreq` and `priority` are always present on every row, and `null` when the sitemap declares nothing for them (most sites publish only some of them; plain-text sitemaps none at all).

```bash
curl -X POST https://api.openscraper.ai/runs \
  -H "Authorization: Bearer sk_live_xxx" \
  -H "Content-Type: application/json" \
  -d '{
    "module": "sitemap_matrix",
    "params": { "url": "https://www.mercedes-benz.com/", "max_results": 5000 },
    "sync": false
  }'
```

A crawl can return hundreds of thousands of URLs, so `result` on the run holds only its first slice. Every row is kept, and the whole set is streamed by an export endpoint — `shape=rows` for one flat object per URL, `shape=report` for the crawler's `{sitemaps_found, pages, errors}` document:

```bash
curl "https://api.openscraper.ai/runs/TASK_ID/export.json?shape=rows" \
  -H "Authorization: Bearer sk_live_xxx"
```

## More modules

Browse every available scraper — with a live, ready-to-run form and copy-paste code snippets — in the [Catalog](https://openscraper.ai/catalog) and the [Playground](https://openscraper.ai/playground). Each module page shows the exact request body for that scraper.
