Sync vs async
Every metered search and scrape endpoint goes through the same job system.wait: trueis the default- a fast terminal result returns
200 - a slower request returns
202withjobId,status,pollUrl, andcreditsCharged
Polling
PollGET /v1/public/jobs/{jobId} with the same bearer token.
202means the job is stillqueuedorrunning200withstatus: "succeeded"includes the terminalresult200withstatus: "failed"includes a terminalerror
completedAtresultTruncatedcached
Webhooks
SendwebhookUrl with any scrape and the worker POSTs the terminal job payload
to that public http(s) URL once the result is persisted. Available to every
organization.
- invalid, private, loopback, and link-local callback URLs are rejected up front
- delivery is best-effort with bounded retries
- a delivery failure never removes the stored result; the job stays pollable
- every attempt sequence is recorded in the delivery log, and can be re-sent
Endpoints and secrets
Each destination is a webhook endpoint with its own signing secret, so the secret you need in order to verify your deliveries cannot forge anybody else’s. Naming awebhookUrl on a job registers that URL if it is not registered
already — the secret exists before the first delivery does. To have it in hand
before you run anything, register up front:
A disabled or deleted endpoint receives nothing. Those attempts are recorded as
skipped rather than dropped, so a gap in your notifications has a visible
cause.
Verifying signatures
Every delivery carries:
To verify: recompute the HMAC over the received timestamp, a
., and the raw
request body; compare with a constant-time compare; and reject deliveries whose
timestamp falls outside a freshness window (±5 minutes recommended). The
timestamp is inside the signed material, which is what stops a captured delivery
being replayed at you later.
X-Scrape-Signature and X-Scrape-Signature-V2 headers are
deprecated and are no longer sent to customer endpoints.
Body shape
status: "failed" and an errorCode — a
caller waiting on a callback should learn about a failure as promptly as a
success, not by timing out.
Deduplicate on jobId. A delivery can arrive more than once (a retry, or a
replay), and the status is terminal, so a repeat carries the same outcome.
The delivery log
GET /v1/public/webhooks/deliveries returns every attempt sequence, newest
first: where it went, what your endpoint answered, and how many attempts it
took. Filter by jobId or status, and page with the returned nextBefore.
This is what makes a missed notification diagnosable rather than a support
thread. Match the X-Scrapebento-Delivery header your receiver logged against
id here.
To re-send one:
The body is rebuilt from the job rather than replayed from a stored copy, so a
replay is exactly what polling the job would return right now. Compare
payloadSha256 in the log to tell whether it differs from the original. If the
job’s result has aged past its retention window, replay returns 410 result_unavailable rather than delivering an empty body and calling it a
success.Idempotency
UseidempotencyKey when a caller might retry the same submission.
- repeated requests with the same key return the same job record
- the worker is not duplicated for a winning request
- billing is intended to follow the same de-duplicated job outcome