SISuperintelligenceDocs

Search docs

Search every page of the documentation.

Self-hosted agents

Job types

The exact input and output of http.batch, browser.batch and exec jobs.

A job's results file is the JSON the job type returns, as below. Types use TypeScript notation; ? marks optional fields.

http.batch

Fetches URLs, one after another. Needs the HTTP requests capability.

Input

{
  requests: {
    id: string                       // your id, echoed in the result
    method: "GET" | "POST"
    url: string                      // http or https
    headers?: Record<string, string>
    body?: string
  }[]
  delayMs?: number                   // wait between requests to a site; default 1000, kept within 250–30,000
}
  • Up to 2,000 requests; more are ignored.
  • Each request may take 30 seconds. Redirects are followed, and the URL a request ends at must be allowed too.
  • Requests identify themselves with User-Agent: si-agent/<version> unless headers sets one.
  • robots.txt is obeyed, with the same rules as browser jobs: the * group and the agent's own group both apply, and a site whose robots.txt can't be read (anything but 200, 404 or 410) is skipped. Requests for /robots.txt itself are always made.
  • delayMs, or the site's crawl-delay when it's longer, is the wait between requests to the same site; a crawl-delay over 60 seconds skips the site.

Output

{
  id: string
  status: number    // the HTTP status, or 0 when there was no response
  body: string      // the response as text; the first 5 MB of characters
  error?: string
}[]                 // in request order
ErrorMeaning
invalid_urlThe URL couldn't be parsed.
invalid_protocolThe URL isn't http or https.
network_deniedThe host isn't on the device's allowlist. The agent files a network request for it, once per host per job.
robots_disallowedrobots.txt doesn't allow the URL for this agent, or the site's robots.txt couldn't be read (anything but 200, 404 or 410).
redirect_deniedA redirect ended on a host outside the allowlist.

Any other failure reports the error's name, such as TimeoutError after 30 seconds or TypeError when the connection fails.

browser.batch

Loads pages in the Chrome, Chromium or Edge installed on the machine, one at a time, and returns data from each. Needs the Browser pages capability. See Browser jobs for how pages are loaded and paced.

Input

{
  pages: {
    id: string                         // 1–200 characters, echoed in the result
    url: string                        // http(s), up to 4,000 characters
    waitForSelector?: string           // a CSS selector to wait for
    waitMs?: number                    // extra wait after load, up to 20,000
    scroll?: boolean                   // scroll down the page first, for lazy-loaded content
    capture?: {
      urlPattern: string               // a JavaScript regular expression matched against response URLs
      max?: number                     // responses to keep; default 10, at most 50
    }
    extract?: ("nextData" | "nuxtData" | "jsonLd" | "html")[]
    select?: {                         // keep only parts of an extract
      nextData?: string | string[]
      nuxtData?: string | string[]
      jsonLd?: string | string[]
    }
  }[]                                  // 1–100 pages
  delayMs?: number                     // between pages; default 8,000, kept within 5,000–60,000
}

extract takes any of these; without it, the defaults are used:

nextData (default), nuxtData (default), jsonLd (default), html

The input is checked when the job is created. A waitForSelector that isn't valid CSS fails the whole job before any page loads.

select shrinks large extracts, such as a __NEXT_DATA__ payload of several megabytes, to what you need. A path is dot-separated keys, array indexes or * for every item: props.pageProps.products.*.name, props.pageProps.total. One path returns its value; several return an object keyed by path. Parts that don't exist are null. For jsonLd, paths start at the array of scripts (*.offers.price).

Output

{
  id: string
  url: string            // as requested
  finalUrl: string       // after redirects
  status: number         // the page's HTTP status, or 0
  captures: { url: string; status: number; body: string }[]
  nextData?: unknown     // JSON of the page's __NEXT_DATA__ script
  nuxtData?: unknown     // the Nuxt payload, decoded to plain JSON
  jsonLd?: unknown[]     // every application/ld+json script, parsed
  html?: string          // the rendered document, up to 2 MB
  error?: string
}[]                      // in page order

A page that timed out after its document arrived still returns what it had, with error: "timeout".

ErrorMeaning
robots_disallowedrobots.txt doesn't allow the page, or the site's robots.txt couldn't be read.
network_deniedThe page's host, or a host it navigated to, isn't on the device's allowlist. The agent files a network request for the host.
timeoutThe page, waitForSelector or extraction didn't finish in time. If the document had arrived, the result still carries what the page had.
navigation_failedThe URL isn't http(s), or the page didn't load.
browser_unavailableNo Chrome, Chromium or Edge was found (or SI_AGENT_CHROME doesn't point to one), or the browser stopped.
output_limitThe job's results reached the size limit, so this page wasn't run.

exec

Runs a shell command on the machine. Needs the Shell commands capability, agents:write on whoever creates the job, and either full network access or a device whose agent limits commands to its allowlist (below).

Input

{
  command: string
  cwd?: string          // working directory; default: the agent's
  timeoutMs?: number    // default 300,000 (5 minutes), at most 1,800,000 (30 minutes)
}

The command runs with /bin/sh -c on macOS and Linux and powershell.exe -NoProfile -Command on Windows, as the user running the agent and with the agent's environment. When the timeout passes, the process is stopped.

Output

{
  code: number | null    // exit code; null when the process was stopped
  stdout: string         // up to 1 MB of characters
  stderr: string         // up to 1 MB of characters
  timedOut: boolean
  canceled?: true        // the job was canceled and the process stopped
  deniedHosts?: string[] // hosts the command tried to reach outside the allowlist
}

Shell commands on an allowlist

On a device with full network access, commands can reach anything. On a device with an allowlist, the agent runs each command in an operating system sandbox that can only connect to a proxy the agent runs; the proxy reaches only the allowed hosts. Tools that use the HTTP_PROXY and HTTPS_PROXY settings (curl, git over https, npm, pip and most HTTP libraries) work for allowed hosts; anything else can't connect. A host outside the allowlist is refused, listed in deniedHosts and filed as a network request.

SystemSandbox
macOSsandbox-exec
LinuxA user and network namespace (unshare), with ip installed. Some distributions turn off unprivileged namespaces.
WindowsNot available: shell commands need full network access.

The agent reports whether its sandbox works when it polls, and only such devices claim exec jobs on an allowlist.