Self-hosted agents
Job types
The exact input and output of http.batch, browser.batch and exec jobs.
A job's results file is the JSON the job type returns, as below. Types use TypeScript notation; ? marks optional fields.
http.batch
Fetches URLs, one after another. Needs the HTTP requests capability.
Input
{
requests: {
id: string // your id, echoed in the result
method: "GET" | "POST"
url: string // http or https
headers?: Record<string, string>
body?: string
}[]
delayMs?: number // wait between requests to a site; default 1000, kept within 250–30,000
}- Up to 2,000 requests; more are ignored.
- Each request may take 30 seconds. Redirects are followed, and the URL a request ends at must be allowed too.
- Requests identify themselves with
User-Agent: si-agent/<version>unlessheaderssets one. - robots.txt is obeyed, with the same rules as browser jobs: the
*group and the agent's own group both apply, and a site whose robots.txt can't be read (anything but 200, 404 or 410) is skipped. Requests for/robots.txtitself are always made. delayMs, or the site's crawl-delay when it's longer, is the wait between requests to the same site; a crawl-delay over 60 seconds skips the site.
Output
{
id: string
status: number // the HTTP status, or 0 when there was no response
body: string // the response as text; the first 5 MB of characters
error?: string
}[] // in request order| Error | Meaning |
|---|---|
invalid_url | The URL couldn't be parsed. |
invalid_protocol | The URL isn't http or https. |
network_denied | The host isn't on the device's allowlist. The agent files a network request for it, once per host per job. |
robots_disallowed | robots.txt doesn't allow the URL for this agent, or the site's robots.txt couldn't be read (anything but 200, 404 or 410). |
redirect_denied | A redirect ended on a host outside the allowlist. |
Any other failure reports the error's name, such as TimeoutError after 30 seconds or TypeError when the connection fails.
browser.batch
Loads pages in the Chrome, Chromium or Edge installed on the machine, one at a time, and returns data from each. Needs the Browser pages capability. See Browser jobs for how pages are loaded and paced.
Input
{
pages: {
id: string // 1–200 characters, echoed in the result
url: string // http(s), up to 4,000 characters
waitForSelector?: string // a CSS selector to wait for
waitMs?: number // extra wait after load, up to 20,000
scroll?: boolean // scroll down the page first, for lazy-loaded content
capture?: {
urlPattern: string // a JavaScript regular expression matched against response URLs
max?: number // responses to keep; default 10, at most 50
}
extract?: ("nextData" | "nuxtData" | "jsonLd" | "html")[]
select?: { // keep only parts of an extract
nextData?: string | string[]
nuxtData?: string | string[]
jsonLd?: string | string[]
}
}[] // 1–100 pages
delayMs?: number // between pages; default 8,000, kept within 5,000–60,000
}extract takes any of these; without it, the defaults are used:
nextData (default), nuxtData (default), jsonLd (default), html
The input is checked when the job is created. A waitForSelector that isn't valid CSS fails the whole job before any page loads.
select shrinks large extracts, such as a __NEXT_DATA__ payload of several megabytes, to what you need. A path is dot-separated keys, array indexes or * for every item: props.pageProps.products.*.name, props.pageProps.total. One path returns its value; several return an object keyed by path. Parts that don't exist are null. For jsonLd, paths start at the array of scripts (*.offers.price).
Output
{
id: string
url: string // as requested
finalUrl: string // after redirects
status: number // the page's HTTP status, or 0
captures: { url: string; status: number; body: string }[]
nextData?: unknown // JSON of the page's __NEXT_DATA__ script
nuxtData?: unknown // the Nuxt payload, decoded to plain JSON
jsonLd?: unknown[] // every application/ld+json script, parsed
html?: string // the rendered document, up to 2 MB
error?: string
}[] // in page orderA page that timed out after its document arrived still returns what it had, with error: "timeout".
| Error | Meaning |
|---|---|
robots_disallowed | robots.txt doesn't allow the page, or the site's robots.txt couldn't be read. |
network_denied | The page's host, or a host it navigated to, isn't on the device's allowlist. The agent files a network request for the host. |
timeout | The page, waitForSelector or extraction didn't finish in time. If the document had arrived, the result still carries what the page had. |
navigation_failed | The URL isn't http(s), or the page didn't load. |
browser_unavailable | No Chrome, Chromium or Edge was found (or SI_AGENT_CHROME doesn't point to one), or the browser stopped. |
output_limit | The job's results reached the size limit, so this page wasn't run. |
exec
Runs a shell command on the machine. Needs the Shell commands capability, agents:write on whoever creates the job, and either full network access or a device whose agent limits commands to its allowlist (below).
Input
{
command: string
cwd?: string // working directory; default: the agent's
timeoutMs?: number // default 300,000 (5 minutes), at most 1,800,000 (30 minutes)
}The command runs with /bin/sh -c on macOS and Linux and powershell.exe -NoProfile -Command on Windows, as the user running the agent and with the agent's environment. When the timeout passes, the process is stopped.
Output
{
code: number | null // exit code; null when the process was stopped
stdout: string // up to 1 MB of characters
stderr: string // up to 1 MB of characters
timedOut: boolean
canceled?: true // the job was canceled and the process stopped
deniedHosts?: string[] // hosts the command tried to reach outside the allowlist
}Shell commands on an allowlist
On a device with full network access, commands can reach anything. On a device with an allowlist, the agent runs each command in an operating system sandbox that can only connect to a proxy the agent runs; the proxy reaches only the allowed hosts. Tools that use the HTTP_PROXY and HTTPS_PROXY settings (curl, git over https, npm, pip and most HTTP libraries) work for allowed hosts; anything else can't connect. A host outside the allowlist is refused, listed in deniedHosts and filed as a network request.
| System | Sandbox |
|---|---|
| macOS | sandbox-exec |
| Linux | A user and network namespace (unshare), with ip installed. Some distributions turn off unprivileged namespaces. |
| Windows | Not available: shell commands need full network access. |
The agent reports whether its sandbox works when it polls, and only such devices claim exec jobs on an allowlist.