Self-hosted agents
Browser jobs
How browser.batch loads pages: which browser, how it identifies itself, robots.txt, pacing and captures.
See Job types for the input and output.
The browser
- The agent uses the Chrome, Chromium or Edge already installed on the machine; it never downloads one. Set
SI_AGENT_CHROMEto the browser's executable to choose one yourself. - Each job gets a fresh browser with a temporary profile, so no cookies or sign-ins carry over between jobs. The browser is closed when the job ends.
- Pages load headless in a 1280 × 900 window, with downloads blocked and the popup blocker on.
Identification
Browser jobs don't disguise themselves:
- The user agent is the browser's own followed by
si-agent/<version>. - The browser's automation flag stays on (
navigator.webdriveristrue). - There are no stealth plugins, fingerprint changes or CAPTCHA solving. A site that challenges automated browsers gets its challenge page back.
robots.txt
Before a site's first page in a job, the agent reads its /robots.txt (just the file: no scripts, and a redirect to a host that isn't allowed doesn't count):
| robots.txt | Result |
|---|---|
200 with rules | Pages must be allowed by the rules. |
404 or 410 | No restrictions. |
Anything else: 401, 403, 5xx, an HTML page, a timeout, no connection | The whole site is skipped for this job. |
How rules apply:
- Rules in the
*group and in a group forsi-agentboth apply: a path must be allowed by each. - Within a group, the longest matching rule wins, and a tie between allow and disallow allows.
*matches any characters and a trailing$anchors the end, as in Google's matcher. Paths include the query string.- Only the first 500 KB of the file is read.
crawl-delay is honored: the longest one from the groups that apply sets the pace. A site asking for more than 60 seconds is skipped.
Disallowed pages return robots_disallowed, and responses from disallowed paths aren't captured.
The allowlist
On a device with an allowlist, every page navigation is checked: the page itself, each redirect, and navigations the page starts. A navigation to a host that isn't allowed is blocked, the page returns network_denied, and the agent files a network request for the host.
The scripts, images and requests a page makes while loading aren't checked, but captures are kept only from allowed hosts. Service workers are bypassed, so they can't answer navigations.
Pacing
Pages load one at a time. Before each page the agent waits the larger of:
delayMs(default 8 seconds, kept between 5 and 60 seconds), and- the site's
crawl-delay.
The wait counts from the previous page, and from the last page of the same site in an earlier job on this device, so back-to-back jobs for one site keep the same pace.
Loading a page
Each page gets 45 seconds, 5 of them reserved for extraction:
- Open the URL and wait for the document.
- Wait for the page to finish loading (up to 15 seconds).
- Wait for
waitForSelector, if given, until the page's time runs out; a missing selector ends the page withtimeout. - Scroll down a screen at a time, if
scrollis set, until the bottom of the page (up to 10 seconds). - Wait
waitMsmore, if given (up to 20 seconds). - Extract the requested data and collect captures.
Captures
With capture, the agent keeps responses the page receives whose URL matches urlPattern, a JavaScript regular expression, for example /api/products\?:
- Only from allowed hosts and paths robots.txt allows.
- Not redirects or
OPTIONSrequests. - Up to
maxresponses (default 10, at most 50), each body cut at 5 MB.
Captures are collected until extraction starts.
Size of the results
A job's results can reach about 50 MB. After that, the remaining pages aren't loaded and return output_limit; queue them in another job.
Try a page
browser-test loads one page through the same code, with full network access and every extraction, and prints what it got. robots.txt still applies.
si-agent browser-test https://www.example.com/products --selector "main" --capture "api/products" --wait 2000 --scrollIt prints the status, the final URL, the size of each extraction and the captured responses.