> ## Documentation Index
> Fetch the complete documentation index at: https://docs.scrapebadger.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Scrape URL

> Scrape a webpage and return its content as HTML, Markdown, or plain text.

## Request Body

<ParamField body="url" type="string" required>
  The URL to scrape. Must be a valid HTTP or HTTPS URL. Private IPs and cloud metadata endpoints are blocked for security.
</ParamField>

<ParamField body="method" type="string" default="GET">
  HTTP method for the request: `GET`, `POST`, `PUT`, `PATCH`, `DELETE`, `HEAD`
  or `OPTIONS`. `POST`, `PUT` and `PATCH` require `request_body`.
</ParamField>

<ParamField body="request_body" type="string">
  Raw body to send with a `POST`/`PUT`/`PATCH`. Pass it as a **string**, and
  set `content_type` to match — there is no `form_data` parameter.

  ```json Submit a form theme={null}
  {
    "url": "https://example.com/login",
    "method": "POST",
    "request_body": "username=alice&password=hunter2",
    "content_type": "application/x-www-form-urlencoded"
  }
  ```

  ```json Call a JSON API theme={null}
  {
    "url": "https://example.com/api/search",
    "method": "POST",
    "request_body": "{\"q\":\"shoes\"}",
    "content_type": "application/json"
  }
  ```
</ParamField>

<ParamField body="content_type" type="string">
  `Content-Type` header for `request_body`. Defaults to
  `application/json` when a body is present.
</ParamField>

<ParamField body="engine" type="string" default="auto">
  Scraping engine tier to use. ScrapeBadger automatically selects the best approach.

  | Value | Engine cost | Description |
  | - | - | - |
  | `auto` | from 1 credit | Automatically picks the best engine for the target site (recommended) |
  | `browser` | 5 credits | Force headless browser with full JavaScript rendering |

  In `auto` mode, simple pages use fast HTTP (1 credit) and JavaScript-heavy
  pages use a browser (5 credits).

  <Warning>
    **The engine cost is not the total.** Every request also pays the
    [`proxy_tier`](#param-proxy-tier) surcharge, so the cheapest possible scrape
    is **2 credits**, not 1, and a browser scrape is **6**, not 5. See
    [Credit costs](#credit-costs) for the full arithmetic.
  </Warning>
</ParamField>

<ParamField body="proxy_tier" type="string" default="simple">
  Proxy quality to route the request through. **This surcharge is added to every
  request**, on top of the engine cost.

  | Value | Pool | Surcharge |
  | - | - | - |
  | `simple` | Datacenter | **+1 credit** |
  | `premium` | Residential | **+8 credits** |
  | `ultra` | Premium residential / mobile | **+8 credits** |

  Stay on `simple` unless the target actually blocks datacenter IPs — moving to
  `premium` makes a browser scrape 13 credits instead of 6. Use
  [`max_cost`](#param-max-cost) if you want a hard ceiling.
</ParamField>

<ParamField body="format" type="string" default="html">
  Output format for the scraped content. Applies to **text** responses only —
  a binary target (image, PDF, archive) ignores it and is returned untouched.

  * `html` — Raw HTML of the page
  * `markdown` — Converted to clean Markdown
  * `text` — Plain text with HTML tags stripped

  These are the only three values. There is no `raw` format — to get an
  unwrapped body, use [`raw_content`](#binary-files-and-raw-bodies).
</ParamField>

<ParamField body="render_js" type="boolean" default={false}>
  Force JavaScript rendering before extracting content. Automatically switches to the `browser` engine. Use this for single-page applications or pages that load content dynamically.
</ParamField>

<ParamField body="wait_for" type="string">
  CSS selector or XPath expression to wait for before extracting content. Only works with the `browser` engine. If `render_js` is `false` and this is set, JS rendering is forced automatically.

  The wait runs **after** `js_scenario`, so a login flow can fill the form,
  submit, and then wait for an element that only exists on the post-login page.
  The wait is a precondition, not a hint: if the selector has not appeared
  within `wait_timeout`, the request fails with `422 wait_for_timeout`, nothing
  is charged, and the response carries `js_scenario_report` so you can see how
  far the flow got. A successful response reports `wait_for_found: true`. For a
  best-effort dwell that never fails, use `wait_after_load` instead.

  ```json theme={null}
  { "wait_for": "#main-content" }
  ```

  ```json theme={null}
  { "wait_for": "//div[@class='results']" }
  ```

  <Warning>
    **`render_js` alone is often not enough.** It returns the DOM once the page
    has loaded — which for anything drawn by a third-party widget (an embedded
    login form, a consent gate, a chat panel, a payment iframe) is the empty
    container, not the content. Those widgets fetch and paint seconds after load.

    If a form or list is missing from your result, name it in `wait_for` rather
    than raising `wait_after_load`; the selector returns as soon as the element
    exists instead of always paying a fixed dwell.

    ```json theme={null}
    {
      "url": "https://example.com/login",
      "render_js": true,
      "wait_for": "#login-container input[type=password]",
      "wait_timeout": 30000
    }
    ```
  </Warning>
</ParamField>

<ParamField body="wait_timeout" type="integer" default={30000}>
  Maximum time in milliseconds to wait for the `wait_for` selector to appear. Range: `1000` – `120000`.
</ParamField>

<ParamField body="wait_after_load" type="integer">
  Additional milliseconds to wait after the page has finished loading, before extracting content. Useful for pages with animations or delayed rendering. Only works with browser engines. Range: `0` – `30000`.
</ParamField>

<ParamField body="js_scenario" type="array">
  A list of browser actions to perform, in order, after the page loads and
  before `wait_for` is awaited and content is extracted. Forces the `browser`
  engine. Each step is an object with an `action` and action-specific
  parameters.

  **Supported actions:**

  | Action | Parameters | Description |
  | - | - | - |
  | `click` | `selector` | Click an element |
  | `fill` | `selector`, `value` | Click into an input and type `value` with real key events |
  | `scroll` | `amount` (optional, px) | Scroll down; a random 300–700 px when `amount` is omitted |
  | `wait` | `timeout` (ms) | Pause for a fixed duration |
  | `evaluate` | `script` | Run JavaScript in the page |

  ```json theme={null}
  {
    "js_scenario": [
      { "action": "fill", "selector": "#username", "value": "me@example.com" },
      { "action": "fill", "selector": "#password", "value": "hunter2" },
      { "action": "click", "selector": "button[type=submit]" },
      { "action": "wait", "timeout": 2000 }
    ]
  }
  ```

  Steps run one after another and stop at the first one that fails (a selector
  that never becomes actionable, a script that throws). A failed step fails the
  request with `422 js_scenario_failed` and nothing is charged. Every response
  that ran a scenario includes `js_scenario_report`, one entry per executed step
  with its outcome and the page URL after it — the log to read when a flow does
  not end where you expected. An unknown action is rejected up front with a
  validation error rather than skipped. (`type` and `milliseconds` are accepted
  as aliases of `action` and `timeout`.)
</ParamField>

<ParamField body="session_id" type="string">
  A unique identifier to persist cookies, fingerprint, and browser storage across multiple requests. Use the same `session_id` to maintain login state or continue a browsing session.

  ```json theme={null}
  { "session_id": "my-shopping-session" }
  ```
</ParamField>

<ParamField body="retry_count" type="integer" default={3}>
  Maximum number of retry attempts when a blocking page is detected. Retries are free — you only pay for the final successful engine. Range: `0` – `10`.
</ParamField>

<ParamField body="retry_on_block" type="boolean" default={true}>
  Whether to automatically retry when a blocking page is detected. Set to `false` to get the blocked response immediately.
</ParamField>

<ParamField body="country" type="string">
  ISO 3166-1 alpha-2 country code for proxy geo-targeting. Routes the request through a proxy in the specified country.

  ```json theme={null}
  { "country": "US" }
  ```

  Common values: `US`, `GB`, `DE`, `FR`, `JP`, `BR`, `AU`.
</ParamField>

<ParamField body="custom_headers" type="object">
  Additional HTTP headers to include in the request to the target URL. Accepts a key-value object.

  ```json theme={null}
  {
    "custom_headers": {
      "Accept-Language": "en-US",
      "Referer": "https://google.com"
    }
  }
  ```
</ParamField>

<ParamField body="eval_js" type="string">
  JavaScript to evaluate in the loaded page after navigation and all waits
  complete. Its return value replaces the page HTML in `content` (objects are
  JSON-serialised). May be an async arrow function. Browser engines only —
  forces the `browser` engine. No extra credits.

  Use it to pull something out of the live page that the HTML does not carry:

  ```json Read the page's cookies theme={null}
  { "url": "https://example.com", "render_js": true, "eval_js": "document.cookie" }
  ```

  ```json Read a value the page computed theme={null}
  {
    "url": "https://example.com",
    "render_js": true,
    "eval_js": "(async () => JSON.stringify({ id: window.__APP_STATE__.id }))()"
  }
  ```

  <Note>
    `document.cookie` returns only non-`HttpOnly` cookies — that is a browser
    rule, not a ScrapeBadger limit. Anti-bot vendors' session cookies are also
    bound to the IP and TLS fingerprint that minted them, so replaying them from
    a different machine does not carry the session. To keep a session alive
    across several scrapes, use [`session_id`](#param-session-id) instead and
    keep the whole flow on our side.
  </Note>
</ParamField>

<ParamField body="screenshot" type="boolean" default={false}>
  Capture a full-page screenshot (PNG). Forces the browser engine. Returned as base64 in the `screenshot_url` response field.
</ParamField>

<ParamField body="video" type="boolean" default={false}>
  Record a video of the browser session (animated GIF). Forces the browser engine. Returned as base64 in the `video_url` response field. No extra charge — you pay the browser engine cost you were already paying. Useful for debugging, visual verification, or monitoring how a page loads.
</ParamField>

<ParamField body="anti_bot" type="boolean" default={false}>
  Attempt to bypass detected anti-bot protection using registered solvers. Adds **+5 credits** to the request cost when a solver is invoked. Only triggered when blocking is actually detected.
</ParamField>

<ParamField body="escalate" type="boolean" default={false}>
  Allow automatic escalation to more powerful engines when the initial engine is blocked.

  Escalation path: `http` → `browser`

  You only pay for the engine that succeeds — costs are **not cumulative**. Without this flag, only the selected engine is tried.
</ParamField>

<ParamField body="max_cost" type="integer">
  Maximum credits to spend on this request. The request fails with a `400` error if the estimated cost would exceed this budget. Useful for controlling costs when using `escalate` or `anti_bot`. Minimum: `1`.
</ParamField>

<ParamField body="raw_content" type="boolean" default={false}>
  Return the body directly as the HTTP response instead of wrapping it in JSON.
  Metadata comes back in `X-Scrape-*` response headers.

  Use it for two things: skipping the JSON encode/decode on large HTML payloads
  (saves 300–1000 ms on 1 MB+ responses), and downloading **binary** files
  without the \~33% base64 overhead. See
  [Binary files and raw bodies](#binary-files-and-raw-bodies).

  Cannot be combined with `ai_extract`, `screenshot` or `video` — those need the
  JSON envelope, so the request falls back to it automatically.
</ParamField>

<ParamField body="ai_extract" type="boolean" default={false}>
  Run AI-powered extraction on the scraped content using the instruction in `ai_prompt`. Adds **+10 credits** to the request cost. The scrape result is still returned even if AI extraction fails, and a failed extraction is not charged.
</ParamField>

<ParamField body="ai_prompt" type="string">
  Natural language instruction for AI data extraction. Required when `ai_extract` is `true`. Maximum 2000 characters.

  ```json theme={null}
  {
    "ai_extract": true,
    "ai_prompt": "Extract all product names and prices as a JSON array"
  }
  ```
</ParamField>

## Credit costs

Every request is billed as **engine + proxy tier + options**. The proxy tier is
charged on every request, which is why nothing costs 1 credit in practice.

| Component | Credits |
| - | - |
| HTTP engine (`http`) | 1 |
| Browser engine (`browser`) | 5 |
| `proxy_tier: simple` (default) | +1 |
| `proxy_tier: premium` / `ultra` | +8 |
| `ai_extract` | +10 |
| `anti_bot`, when a solver is actually invoked | +5 |
| `screenshot`, `video`, `js_scenario`, `eval_js`, `wait_for` | +0 |

Common totals:

| Request | Total |
| - | - |
| Simple page, default tier | **2** |
| JS-rendered page, default tier | **6** |
| JS-rendered page, `premium` tier | **13** |
| Simple page + `ai_extract` | **12** |

The exact amount charged is always in the `X-Credits-Used` response header and
the `credits_used` body field — trust those over any estimate. A request
rejected before scraping (`400`, `422`) is charged **0**.

<Tip>
  Set [`max_cost`](#param-max-cost) to put a hard ceiling on a request. It is
  checked against the estimate *including* the proxy surcharge.
</Tip>

## Response

<ResponseField name="success" type="boolean">
  Whether the scrape completed successfully. `false` when all retries are exhausted and the page is still blocked.
</ResponseField>

<ResponseField name="url" type="string">
  The final URL after any redirects.
</ResponseField>

<ResponseField name="status_code" type="integer">
  HTTP status code from the target URL.
</ResponseField>

<ResponseField name="content" type="string">
  The scraped content in the requested format. `null` when `success` is `false`,
  and `null` for binary targets — those come back in `content_base64`.
</ResponseField>

<ResponseField name="content_base64" type="string">
  Base64-encoded response body, returned instead of `content` when the target
  serves a binary payload. Only present when `is_binary` is `true`. Bodies above
  25 MB are not base64-encoded — re-request those with `raw_content: true`.
</ResponseField>

<ResponseField name="is_binary" type="boolean">
  Whether the target returned a binary (non-text) body.
</ResponseField>

<ResponseField name="content_type" type="string">
  The target's response `Content-Type`, normalised to the bare media type
  (e.g. `image/jpeg`).
</ResponseField>

<ResponseField name="format" type="string">
  The output format used: `html`, `markdown`, or `text`.
</ResponseField>

<ResponseField name="engine_used" type="string">
  The engine tier that produced the final result.
</ResponseField>

<ResponseField name="credits_used" type="integer">
  Total credits charged for this request: engine cost + `proxy_tier` surcharge,
  plus solver and AI extraction when used. See [Credit costs](#credit-costs).
  This is the authoritative figure and matches the `X-Credits-Used` header.
</ResponseField>

<ResponseField name="duration_ms" type="integer">
  Total request processing time in milliseconds.
</ResponseField>

<ResponseField name="retries_used" type="integer">
  Number of retry attempts performed. `0` if the first attempt succeeded.
</ResponseField>

<ResponseField name="content_length" type="integer">
  Size of the returned content in bytes.
</ResponseField>

<ResponseField name="screenshot_url" type="string">
  Base64-encoded PNG screenshot of the page. Only present when `screenshot: true` was requested.
</ResponseField>

<ResponseField name="video_url" type="string">
  Base64-encoded animated GIF of the browser session. Only present when `video: true` was requested.
</ResponseField>

<ResponseField name="headers" type="object">
  HTTP response headers from the target URL.
</ResponseField>

<ResponseField name="blocking_detected" type="boolean">
  Whether a blocking page was detected during scraping.
</ResponseField>

<ResponseField name="blocking_details" type="object">
  Details about the detected blocking page. Only present when `blocking_detected` is `true`.

  <Expandable>
    <ResponseField name="is_blocked" type="boolean">
      Whether the page is confirmed as a blocking page.
    </ResponseField>

    <ResponseField name="block_type" type="string">
      Type of block detected (e.g., `cloudflare`, `datadome`, `akamai`, `kasada`).
    </ResponseField>

    <ResponseField name="confidence" type="number">
      Confidence score from `0.0` to `1.0`.
    </ResponseField>

    <ResponseField name="details" type="string">
      Human-readable description of the block.
    </ResponseField>
  </Expandable>
</ResponseField>

<ResponseField name="antibot_systems" type="array">
  List of anti-bot systems detected on the page.

  <Expandable>
    <ResponseField name="system" type="string">
      System name (e.g., `cloudflare_turnstile`, `datadome`, `akamai`, `kasada`, `amazon_waf`).
    </ResponseField>

    <ResponseField name="confidence" type="number">
      Confidence score from `0.0` to `1.0`.
    </ResponseField>

    <ResponseField name="details" type="string">
      Additional detection details.
    </ResponseField>
  </Expandable>
</ResponseField>

<ResponseField name="captcha_systems" type="array">
  List of CAPTCHA systems detected on the page.

  <Expandable>
    <ResponseField name="system" type="string">
      System name (e.g., `recaptcha_v2`, `recaptcha_v3`, `hcaptcha`, `geetest`).
    </ResponseField>

    <ResponseField name="confidence" type="number">
      Confidence score from `0.0` to `1.0`.
    </ResponseField>

    <ResponseField name="details" type="string">
      Additional detection details.
    </ResponseField>
  </Expandable>
</ResponseField>

<ResponseField name="anti_bot_solved" type="boolean">
  Whether the anti-bot solver successfully bypassed the protection.
</ResponseField>

<ResponseField name="solver_used" type="string">
  Name of the solver that successfully bypassed the block. `null` if no solver was used.
</ResponseField>

<ResponseField name="wait_for_found" type="boolean">
  Whether the `wait_for` selector appeared within `wait_timeout`. `null` when no
  `wait_for` was requested. A miss never reaches you as a `200` — it is a
  `422 wait_for_timeout`.
</ResponseField>

<ResponseField name="js_scenario_report" type="array">
  One entry per executed `js_scenario` step, in order. `null` when no scenario
  ran. Execution stops at the first failed step, so a report shorter than the
  scenario means a step broke.

  <Expandable>
    <ResponseField name="step" type="integer">
      Zero-based index of the step in your `js_scenario`.
    </ResponseField>

    <ResponseField name="action" type="string">
      The step's action, as you sent it.
    </ResponseField>

    <ResponseField name="selector" type="string">
      The step's selector, if it had one.
    </ResponseField>

    <ResponseField name="ok" type="boolean">
      Whether the step completed.
    </ResponseField>

    <ResponseField name="error" type="string">
      Why the step failed (e.g. `TimeoutError: waiting for locator("#go")`). `null` on success.
    </ResponseField>

    <ResponseField name="url" type="string">
      The page URL right after the step — the quickest way to see whether a submit actually navigated.
    </ResponseField>
  </Expandable>
</ResponseField>

<ResponseField name="ai_extraction" type="object | string | array">
  Structured data extracted by the LLM based on `ai_prompt`. The shape depends on your prompt. `null` when `ai_extract` is `false` or extraction failed.
</ResponseField>

<ResponseField name="ai_model" type="string">
  The LLM model used for extraction (e.g., `gpt-4o-mini`). `null` when AI extraction was not used.
</ResponseField>

<ResponseField name="ai_error" type="string">
  Error message if AI extraction failed. The scrape result is still returned. `null` on success.
</ResponseField>

## Binary files and raw bodies

`/v1/web/scrape` handles binary targets — images, PDFs, archives, fonts, audio
and video — as well as HTML. The bytes are returned exactly as the origin sent
them; nothing is decoded, parsed or re-encoded on the way through.

This works through the same anti-bot machinery as a page scrape, so a file
behind Cloudflare, DataDome or Imperva is fetched with the same engine, proxy
tier and session you would use for its parent page. Pass the same `session_id`
you used to scrape the page the file was linked from and the file download
reuses that session's cookies and fingerprint.

There are two ways to get the bytes.

**`raw_content: true` — the body itself.** No base64 overhead. Best for large
files and for piping straight to disk.

<CodeGroup>
  ```bash cURL theme={null}
  curl -X POST "https://scrapebadger.com/v1/web/scrape" \
    -H "X-API-Key: YOUR_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "url": "https://example.com/photos/item.jpg",
      "raw_content": true,
      "session_id": "gallery-session"
    }' \
    --output item.jpg
  ```

  ```python Python theme={null}
  import requests

  response = requests.post(
      "https://scrapebadger.com/v1/web/scrape",
      headers={"X-API-Key": "YOUR_API_KEY"},
      json={
          "url": "https://example.com/photos/item.jpg",
          "raw_content": True,
          "session_id": "gallery-session",
      },
  )

  with open("item.jpg", "wb") as f:
      f.write(response.content)          # exact bytes — do NOT use response.text

  print(response.headers["Content-Type"])
  print(response.headers["X-Credits-Used"])
  ```

  ```javascript JavaScript theme={null}
  const response = await fetch("https://scrapebadger.com/v1/web/scrape", {
    method: "POST",
    headers: {
      "X-API-Key": "YOUR_API_KEY",
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      url: "https://example.com/photos/item.jpg",
      raw_content: true,
      session_id: "gallery-session",
    }),
  });

  const buffer = Buffer.from(await response.arrayBuffer());
  ```
</CodeGroup>

<Warning>
  Read the response as **bytes**, not text. `response.text` in Python or
  `response.text()` in JavaScript will decode the payload as a string and
  corrupt it. Use `response.content` / `response.arrayBuffer()`.
</Warning>

**JSON with `content_base64` — one request, metadata included.** Use this when
you also want `credits_used`, `engine_used` or the protection detections
alongside the file.

```python Python theme={null}
import base64
import requests

data = requests.post(
    "https://scrapebadger.com/v1/web/scrape",
    headers={"X-API-Key": "YOUR_API_KEY"},
    json={"url": "https://example.com/report.pdf"},
).json()

if data["is_binary"]:
    with open("report.pdf", "wb") as f:
        f.write(base64.b64decode(data["content_base64"]))
    print(data["content_type"])   # application/pdf
else:
    print(data["content"])        # a text target — HTML/markdown/text
```

Binary responses served with `raw_content` always carry
`Content-Disposition: attachment` and `X-Content-Type-Options: nosniff`, and
their `Content-Type` is restricted to a known-safe set — scraped bytes are
never labelled in a way that would let a browser execute them.

Bodies larger than 25 MB are not base64-encoded into JSON. Those return
`content_base64: null` with a `detail` telling you to re-request using
`raw_content: true`, which returns them without the base64 expansion. Very
large files are still held in memory end to end, so treat a few hundred MB as
the practical ceiling for a single request.

## Examples

### Basic scrape

<CodeGroup>
  ```bash cURL theme={null}
  curl -X POST "https://scrapebadger.com/v1/web/scrape" \
    -H "x-api-key: YOUR_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{"url": "https://scrapebadger.com", "format": "markdown"}'
  ```

  ```python Python theme={null}
  import requests

  response = requests.post(
      "https://scrapebadger.com/v1/web/scrape",
      headers={"x-api-key": "YOUR_API_KEY"},
      json={"url": "https://scrapebadger.com", "format": "markdown"}
  )
  print(response.json()["content"])
  ```

  ```javascript JavaScript theme={null}
  const res = await fetch("https://scrapebadger.com/v1/web/scrape", {
    method: "POST",
    headers: { "x-api-key": "YOUR_API_KEY", "Content-Type": "application/json" },
    body: JSON.stringify({ url: "https://scrapebadger.com", format: "markdown" })
  });
  const data = await res.json();
  console.log(data.content);
  ```
</CodeGroup>

### JavaScript rendering with wait

```json theme={null}
{
  "url": "https://scrapebadger.com/spa-page",
  "format": "html",
  "render_js": true,
  "wait_for": "#dynamic-content",
  "wait_timeout": 10000
}
```

### AI extraction

```json theme={null}
{
  "url": "https://scrapebadger.com/products",
  "format": "markdown",
  "ai_extract": true,
  "ai_prompt": "Extract all product names, prices, and ratings as a JSON array of objects with keys: name, price, rating"
}
```

### Full anti-bot bypass with budget

```json theme={null}
{
  "url": "https://heavily-protected-site.com",
  "format": "markdown",
  "escalate": true,
  "anti_bot": true,
  "max_cost": 20,
  "country": "US"
}
```

### Browser automation scenario

Log in, then wait for an element that only exists once the login succeeded.
If the flow never gets there you receive a free `422` with the step-by-step
report instead of a billed copy of the login page.

```json theme={null}
{
  "url": "https://example.com/login",
  "js_scenario": [
    { "action": "fill", "selector": "#username", "value": "me@example.com" },
    { "action": "fill", "selector": "#password", "value": "hunter2" },
    { "action": "click", "selector": "button[type=submit]" }
  ],
  "wait_for": "#account-dashboard",
  "wait_timeout": 60000
}
```

Infinite scroll:

```json theme={null}
{
  "url": "https://scrapebadger.com/infinite-scroll",
  "format": "text",
  "js_scenario": [
    { "action": "scroll", "amount": 1000 },
    { "action": "wait", "timeout": 2000 },
    { "action": "scroll", "amount": 1000 },
    { "action": "wait", "timeout": 2000 }
  ]
}
```

## Error Responses

| Status | Description |
| - | - |
| `400` | Invalid URL, cost exceeds `max_cost`, or requested engine not available |
| `402` | Insufficient credits |
| `422` | Blocking detected after all retries exhausted (`success: false`, `blocking_details` populated); the `wait_for` selector never appeared (`error: "wait_for_timeout"`); a `js_scenario` step failed (`error: "js_scenario_failed"`); or the body contained an unknown field (`error: "unknown_request_field"`). Never charged. |
| `429` | Rate limit exceeded |
| `500` | Unexpected server error |

### Unknown request fields

A field this endpoint does not define is rejected rather than ignored, so a
request never looks like it succeeded at something it did not do. The response
names the offending field and lists everything the endpoint accepts, and
nothing is charged.

```json 422 — Unknown request field theme={null}
{
  "error": "unknown_request_field",
  "detail": "Unknown request field(s): return_cookies. These were silently ignored by older versions; they are now rejected so a request never appears to succeed at something it did not do.",
  "unknown_fields": ["return_cookies"],
  "accepted_fields": ["ai_extract", "ai_prompt", "anti_bot", "..."]
}
```

There is no `return_cookies` or `form_data` parameter. To read cookies use
[`eval_js`](#param-eval-js); to send a request body use
[`method`](#param-method) with `request_body`.

<ResponseExample>
  ```json 200 — Success theme={null}
  {
    "success": true,
    "url": "https://scrapebadger.com",
    "status_code": 200,
    "content": "# Example Domain\n\nThis domain is for use in illustrative examples...",
    "format": "markdown",
    "engine_used": "http",
    "credits_used": 2,
    "duration_ms": 342,
    "retries_used": 0,
    "content_length": 1256,
    "screenshot_url": null,
    "video_url": null,
    "headers": {
      "content-type": "text/html; charset=UTF-8"
    },
    "blocking_detected": false,
    "blocking_details": null,
    "antibot_systems": [],
    "captcha_systems": [],
    "anti_bot_solved": false,
    "solver_used": null,
    "ai_extraction": null,
    "ai_model": null,
    "ai_error": null
  }
  ```

  ```json 200 — With AI Extraction theme={null}
  {
    "success": true,
    "url": "https://scrapebadger.com/products",
    "status_code": 200,
    "content": "# Products\n\n...",
    "format": "markdown",
    "engine_used": "http",
    "credits_used": 12,
    "duration_ms": 2150,
    "retries_used": 0,
    "content_length": 8432,
    "screenshot_url": null,
    "video_url": null,
    "headers": {},
    "blocking_detected": false,
    "blocking_details": null,
    "antibot_systems": [],
    "captcha_systems": [],
    "anti_bot_solved": false,
    "solver_used": null,
    "ai_extraction": [
      { "name": "Widget Pro", "price": "$29.99", "rating": 4.5 },
      { "name": "Widget Basic", "price": "$9.99", "rating": 4.2 }
    ],
    "ai_model": "gpt-4o-mini",
    "ai_error": null
  }
  ```

  ```json 422 — Blocked theme={null}
  {
    "success": false,
    "url": "https://protected-site.com",
    "status_code": 403,
    "content": null,
    "format": "markdown",
    "engine_used": "http",
    "credits_used": 0,
    "duration_ms": 5230,
    "retries_used": 3,
    "content_length": 0,
    "screenshot_url": null,
    "video_url": null,
    "headers": {},
    "blocking_detected": true,
    "blocking_details": {
      "is_blocked": true,
      "block_type": "cloudflare",
      "confidence": 0.95,
      "details": "Cloudflare challenge page detected"
    },
    "antibot_systems": [
      { "system": "cloudflare_turnstile", "confidence": 0.95, "details": null }
    ],
    "captcha_systems": [],
    "anti_bot_solved": false,
    "solver_used": null,
    "ai_extraction": null,
    "ai_model": null,
    "ai_error": null
  }
  ```
</ResponseExample>


## OpenAPI

````yaml POST /v1/web/scrape
openapi: 3.1.0
info:
  title: ScrapeBadger Web Scraping API
  version: 1.0.0
  description: Web scraping API with anti-bot bypass, JS rendering, and AI extraction.
servers:
  - url: https://scrapebadger.com
    description: Production
security:
  - apiKeyAuth: []
paths:
  /v1/web/scrape:
    post:
      summary: Scrape URL
      description: >-
        Scrape a webpage and return its content as HTML, Markdown, or plain
        text.
      operationId: scrapeUrl
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required:
                - url
              properties:
                url:
                  type: string
                  description: The URL to scrape. Must be a valid HTTP or HTTPS URL.
                method:
                  type: string
                  default: GET
                  enum:
                    - GET
                    - POST
                    - PUT
                    - PATCH
                    - DELETE
                    - HEAD
                    - OPTIONS
                  description: HTTP method. POST/PUT/PATCH require request_body.
                request_body:
                  type: string
                  description: >-
                    Raw body for POST/PUT/PATCH, as a string. There is no
                    form_data parameter — set content_type to match.
                content_type:
                  type: string
                  description: >-
                    Content-Type for request_body. Defaults to application/json
                    when a body is present.
                proxy_tier:
                  type: string
                  default: simple
                  enum:
                    - simple
                    - premium
                    - ultra
                  description: >-
                    Proxy pool to route through. Its surcharge is added to EVERY
                    request: simple +1, premium +8, ultra +8 credits.
                eval_js:
                  type: string
                  description: >-
                    JavaScript evaluated in the loaded page after navigation and
                    waits; its return value replaces the page HTML in content.
                    Browser engines only. Use it to read what the HTML does not
                    carry, e.g. document.cookie (non-HttpOnly cookies only, and
                    anti-bot cookies are bound to the IP and TLS fingerprint
                    that minted them).
                engine:
                  type: string
                  default: auto
                  enum:
                    - auto
                    - browser
                  description: Scraping engine tier to use.
                format:
                  type: string
                  default: html
                  enum:
                    - html
                    - markdown
                    - text
                  description: Output format for the scraped content.
                render_js:
                  type: boolean
                  default: false
                  description: Force JavaScript rendering.
                wait_for:
                  type: string
                  description: CSS selector or XPath to wait for before extracting.
                wait_timeout:
                  type: integer
                  default: 30000
                  description: Max wait time in ms for wait_for selector.
                wait_after_load:
                  type: integer
                  description: Additional ms to wait after page load.
                js_scenario:
                  type: array
                  items:
                    type: object
                  description: Browser actions to perform before extracting.
                session_id:
                  type: string
                  description: Persist cookies and state across requests.
                retry_count:
                  type: integer
                  default: 3
                  description: Max retry attempts on blocking detection.
                retry_on_block:
                  type: boolean
                  default: true
                  description: Auto-retry on blocking page detection.
                country:
                  type: string
                  description: ISO 3166-1 alpha-2 country code for proxy geo-targeting.
                custom_headers:
                  type: object
                  description: Additional HTTP headers for the target request.
                screenshot:
                  type: boolean
                  default: false
                  description: Capture a full-page PNG screenshot.
                video:
                  type: boolean
                  default: false
                  description: Record browser session as animated GIF. No extra charge.
                anti_bot:
                  type: boolean
                  default: false
                  description: Attempt anti-bot bypass when blocking detected.
                escalate:
                  type: boolean
                  default: false
                  description: Allow auto-escalation to stronger engines.
                max_cost:
                  type: integer
                  description: Maximum credits budget for this request.
                ai_extract:
                  type: boolean
                  default: false
                  description: Run AI extraction on scraped content.
                ai_prompt:
                  type: string
                  description: Natural language instruction for AI extraction.
      responses:
        '200':
          description: Successful scrape
          content:
            application/json:
              schema:
                type: object
                properties:
                  success:
                    type: boolean
                  url:
                    type: string
                  status_code:
                    type: integer
                  content:
                    type: string
                  format:
                    type: string
                  engine_used:
                    type: string
                  credits_used:
                    type: integer
                  duration_ms:
                    type: integer
                  retries_used:
                    type: integer
                  content_length:
                    type: integer
                  screenshot_url:
                    type: string
                    nullable: true
                  video_url:
                    type: string
                    nullable: true
                  headers:
                    type: object
                  blocking_detected:
                    type: boolean
                  blocking_details:
                    type: object
                    nullable: true
                  antibot_systems:
                    type: array
                  captcha_systems:
                    type: array
                  anti_bot_solved:
                    type: boolean
                  solver_used:
                    type: string
                    nullable: true
                  ai_extraction:
                    nullable: true
                  ai_model:
                    type: string
                    nullable: true
                  ai_error:
                    type: string
                    nullable: true
components:
  securitySchemes:
    apiKeyAuth:
      type: apiKey
      in: header
      name: x-api-key

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.