Written to be checkable. Where something is not in place, this page says so rather than leaving it out, because a security page you cannot trust on the gaps is not worth trusting on the rest either.

API keys

Keys are 32 bytes from the operating system’s CSPRNG, base64url-encoded behind a satk_ prefix. We cannot read your key. Only a SHA-256 digest is stored, alongside the last four characters so you can tell two keys apart in the dashboard. A key is shown once, at creation. If you lose it, the only remedy is to revoke it and create another — we have nothing to send you. If a key does leak, revoke it first and read GET /v1/public/requests afterwards — it lists every job the key ran, so the blast radius is a query rather than a guess.

Webhooks

Each registered endpoint gets its own signing secret. A shared secret would mean any customer who could read theirs could forge deliveries to anyone else’s endpoint; a per-endpoint secret makes that impossible by construction. Every delivery carries: Verify the signature and the timestamp:
Every attempt — delivered or not — is recorded and can be re-sent from GET /v1/public/webhooks/deliveries. The payload body is deliberately not stored. A delivery log holding every result would be a second copy of your scraped data, kept for longer, in a place designed for reading. Secrets are fetched one at a time from GET /v1/public/webhooks/endpoints/{id}/secret rather than appearing in list responses — a secret in every list response ends up in far more logs and screenshots than one you have to ask for.

Your data

The research corpus is per-organization: retrieval spans your prior runs and never anyone else’s. That isolation is what makes topic memory a feature rather than a leak. Scraped content is fetched on your behalf from public sources and stored to serve your job and your cache window. We do not train models on it, and it is never shared between organizations.

Content sent to third parties

Some endpoints cannot do their job locally, and it would be misleading to leave that out of a page about where your data goes. Everything else — url, search, seo, infra, rss, sitemap, mentions, hackernews/*, linkedin/employees — is processed entirely on our own infrastructure and reaches no model provider. Which provider serves those actions is an operator configuration, and provider terms differ on whether API traffic may be retained or used for training. If that distinction matters to you, ask us which providers are configured before you send anything sensitive through the four endpoints above. Proxies see the destination URL and the fetched bytes for every scrape, because that is what a proxy does.

Accounts

Identity is Ory Kratos. Passwords are never handled by our application code.
  • Two-factor: TOTP, opt-in per account. Passkeys are scaffolded but off by default — do not count on them yet.
  • Sessions: HttpOnly, SameSite=Lax, Secure cookies, scoped to an explicit domain rather than one derived from the request. User and admin sessions use different cookie names, so a valid user session cannot be replayed against an admin endpoint.
  • Step-up: destructive account operations require a second factor when one is enrolled.
  • Sessions last 7 idle days.

Rate limits and isolation

Per-token limits are enforced in shared state, so they hold across every API replica rather than per-process. Queued and running jobs are capped per organization: one tenant’s burst cannot starve another’s throughput. Requests to internal and private network ranges are refused, so a URL you submit cannot be used to reach our infrastructure.

What we do not have

The honest part.
  • No SOC 2 report. Type I is planned, not started. If your procurement needs one today, we are not a fit yet.
  • No third-party penetration test. Not yet commissioned.
  • No SSO or SCIM. Planned; not built.
  • No contractual uptime SLA. We publish per-endpoint reliability with stated targets and say when we miss them, which is a different and more checkable thing than a credit-backed promise.
  • No data-residency choice. Everything runs in one region.

Reporting a vulnerability

Email security@scrapebento.com. Please do not open a public issue. We aim to acknowledge within one business day. We will not pursue legal action for good-faith research that stays within your own organization’s data, avoids degrading the service for others, and gives us a reasonable chance to fix the issue before disclosure.