Limits are per API token, applied after authentication — so they cannot be dodged by changing IP, and one token’s traffic never eats another’s budget.

Two independent buckets

They are separate on purpose. Polling a long-running research job cannot consume the budget you need to submit work, so a job that takes minutes does not throttle the rest of your pipeline. /v1/public/balance, /usage, and /requests share the poll bucket. Checking your balance never costs you create capacity. The webhook registry and delivery log share it too; replaying a delivery sends a real request, so that one is on the create bucket. Your bucket is shared across every server that answers your requests, so the number above is the number you get — not that number multiplied by however many machines happen to be running.

Headers

Every response carries your current position: Read X-RateLimit-Remaining as you go rather than waiting for a 429 — it lets you pace a batch instead of recovering from a rejection.

Handling 429

Wait Retry-After seconds, then retry. Exponential backoff on top is sensible if you are running many workers, since they will all be told the same number and would otherwise retry in lockstep.
A 429 is never charged.

Concurrency, not just rate

Long-running kinds hold a worker slot for their whole run — research and company up to 10 minutes, ads 8, linkedin_employees 6. Submitting a thousand of them inside your rate limit is possible and will not make them finish sooner; they queue. For bulk work, prefer POST /v1/public/scrape/batch, which is designed for it, and use webhookUrl or polling rather than holding connections open.

If you need more

The defaults suit an interactive integration. If you are running a sustained pipeline, get in touch before working around the limit — raising it for your token is usually the right answer and is easier than an elaborate retry scheme.

Per-call ceilings

Rate limits bound how many calls you may make. A separate ceiling bounds how much one call may ask for, because a request for a thousand items is a different amount of work from a request for fifty and costs the same credits. A limit above the ceiling is clamped, not refused — you get the ceiling and a successful response, never a 400 for asking too much. The same applies to deep on /scrape/channel: without the extended ceiling the flag is ignored and you get a bounded page. Several of these endpoints return a nextCursor and an exhausted flag; pass the cursor back to collect the next page. Where an endpoint has no cursor, one call is one page of whatever the target publishes. The extended column is enabled per organization rather than by plan. If you are backfilling history rather than polling for new items, get in touch — that is what it is for, and it is cheaper for both of us than working around the ceiling.