Web scraper for agents: blocked, empty and wrong pages reported as such, with an Evidence Record.
A transparent, verifiable web extraction system built for RAG and Agent workflows.
Try it without installing anything at octocrawl.dev, read the documentation, or connect your agent to hosted Octocrawl in one line, no key needed to start (how, for Claude Code, Cursor, OpenCode and Codex):
claude mcp add --transport http octocrawl https://mcp.octocrawl.dev/mcp
Most crawlers report "success" when they return empty pages, challenge screens, or the wrong content. Octocrawl makes failure visible and fixable:
Before (typical crawler):
✓ Fetched example.com/article
Status: 200 OK
Content: 953 bytes
After (Octocrawl):
✗ Fetched example.com/article
Status: blocked (cloudflare_challenge)
Lane: http → escalated to browser_local
Evidence: artifacts=[] (no screenshot or DOM snapshot was produced)
Cost: 847 tokens, 2.3s, $0.0042
Fix: needs user login or proxy (tier 1b/2)
empty_verified, blocked, failed with reasons, not silent empties; a page answered with an error status keeps its httpStatus and Markdown as evidence, never as success, and so does a page on which Octocrawl finds no main content (failed/empty_unverified with the whole page's Markdown), an article region that only names the page included (headings of fewer than 100 characters and almost nothing beside them, no image the Markdown carries, in-page jump links such as "Skip to Filters" set aside)artifacts: [] is an explicit empty artifact list, not a promise that every failed page has a screenshot or DOM snapshot; browser bytesWire: null means wire bytes were not measuredevidenceRecord (final URL, redirect chain, fetch time, status and reason, lane, robots.txt decision, raw and output hashes, field evidence), stated the same way in every lane and described by a versioned JSON Schema; see the referenceThe no-install, single-page web preview runs at octocrawl.dev
(five previews a day; see Public preview for how it is deployed). Besides Markdown it
returns a page's links and metadata, and up to 20 fields read from the page without a model. The same site serves
the documentation: MCP
connection steps for four clients, four task guides, and limits. The pages are generated from
Markdown in this repository during the public web build;
npm run public:preview:local serves them at http://127.0.0.1:8798/docs/.
Install nothing first: with Node.js 22.13+ or 24+,
npx octocrawl scrape https://example.com --markdown
pip install octocrawl-client # the Python client of a running API (npx octocrawl serve)
npm install @octocrawl/sdk # the TypeScript client; @octocrawl/mcp is the MCP server
The packages are published from this repository (0.3.0 on 2026-10-05); the Install check workflow runs the npx line from an empty cache on macOS, Windows and Linux each week. To work on the code, use Node.js 22.13+ or 24+ and npm (the PDF text engine, pdf.js, needs 22.13 or later) and build from this checkout. For the full Monitor → result → HTTPS event → restart workflow, follow the onboarding guide and independent developer acceptance checklist.
git clone https://github.com/77777R7/w2l.git
cd w2l
npm ci
npx playwright install chromium
npm run typecheck
npm test
npm run scrape -- https://example.com --markdown
npm run crawl -- https://example.com --max-pages 20
octocrawl (the @octocrawl/cli package, also published as octocrawl; npm run w2l -- <command> in a checkout) runs the API's engine in its own process, so every option of the REST API works on the command line. Each option is a flag under its kebab-case name: maxAge is --max-age, onlyMainContent: false is --no-only-main-content, includeTags takes a,b and may repeat, formats takes markdown,tables or a JSON array for an entry with options, parsers takes pdf, none or JSON, headers takes JSON or --header name=value (repeatable), and the URLs are arguments. The request is checked by the REST API's own parser, so a value the API refuses is refused here with the same message (exit code 2).
octocrawl scrape <url> prints the scrape response as JSON (compact, as MCP gets it; --debug for the full one), or the Markdown alone with --markdown.octocrawl batch <url>... (or --urls-file <file>, one URL per line) and octocrawl crawl <url> run to the end and print { report, items }. Ctrl-C leaves the job paused: octocrawl crawl --resume <taskId> continues a crawl (a crawl recorded as pending or running is refused, as another process may be running it), and a batch resumes when the API starts on the same task root. --webhook is not offered, since a command runs no delivery worker: send the job to octocrawl serve instead.octocrawl map <url> prints the map.--out <dir> writes results.jsonl (one result per line), results.csv (one row per page with its evidence, failed pages included: the columns of the Python client's to_pandas() without the Markdown, then markdown_file), report.json for a job, and per page <n>-<host-path>.md with its Markdown and <n>-<host-path>.table-<i>.csv for each table of the tables format.octocrawl serve runs the local API, as npm run api does, with the same flags (--port, --host, --hosted, --token, ...).octocrawl login import <site> saves your login to a site from the Chrome you already use, so --mode authed reads its pages signed in as you; octocrawl login list and octocrawl login remove <site> show and forget saved logins without printing a cookie. See Your own login (mode authed).Exit codes: 0 for a page read as content or a completed job, 1 for anything else Octocrawl answered, 2 for a refused command line, 130 when interrupted. The task root, where tasks, saved files and the page cache live, is --task-root, else W2L_TASK_ROOT, else .w2l/cli, apart from the API's .w2l/api; never point a command at the task root of a running API server, which could run the same job twice. A command never resumes the task root's earlier jobs, as the API does when it starts. The earlier in-process ladder tool is w2l-ladder (and w2l-fetch) in @w2l/bench.
Mode authed reads a page with the login you saved for its site, in Octocrawl's own browser. To save one, sign in to the site in Google Chrome (144 or later), in the default profile and a normal window; open chrome://inspect/#remote-debugging and turn on "Allow remote debugging for this browser instance" once; then run octocrawl login import example.com (a domain or a page URL). Chrome asks "Allow remote debugging?" for every connection: click Allow. Octocrawl connects once, reads that site's cookies (its own, a parent domain's, and its subdomains' when the site sets cookies on the name itself) and the localStorage of the site's tabs you have open, saves them and disconnects; it never reads Chrome's files on disk, and reads a tab's storage without loading anything or running script in it. Logins are kept in W2L_SESSIONS_FILE, else ~/.w2l/sessions.json, readable by you alone, and the command line, the local API and the local MCP service read the same file. Records carry the login's SHA-256, cookie count and localStorage origins and item count, never a value. The same import works without the command line on a server running on your machine: POST /v1/logins/import with { "site": "example.com" } (and approveTimeoutMs, how long to wait for your Allow, 10 s to 10 min, default 2 min) answers { domain, savedAt, cookieCount, localStorage, localStorageRead, sessionSha256 } once you click Allow (localStorage: { origins, itemCount } saved, or null; localStorageRead false when no tab of the site was open, so none was read; localStorageUnread, the origins of open tabs Chrome did not give the storage of, a crashed or discarded tab, which a reload brings back). GET /v1/logins lists the saved logins and DELETE /v1/logins/:site forgets one. The SDK's importLogin, listLogins and removeLogin, and the local MCP tools import_login, list_logins and remove_login, call these routes. A local MCP server offers all three. import_login lets an agent choose which site's login to save, so its description tells the agent to ask you first, and Chrome's Allow is still yours to click. No route answers a cookie. A server that is not on loopback, the hosted service, and a server without your Chrome refuse an import with 409.
example.com covers www.example.com and other subdomains. When one applies, the authed rung goes first, before the public rungs, which would take a logged-out page as the answer; if the site refuses the login (login_wall, for example when it expired), the run ends there instead of returning the logged-out page. A refusal is a login_wall page, a redirect to the site's login page, or a page that asks for a sign-in in place: a heading or plain line among its first eight that begins with the request, such as "Log in to view your wishlists", "Please sign in", "You must be logged in" or "Login required". A header's Sign in link, a table row, a list item, a line with a link, a line where the words come after others ("Step 2: Sign in to continue") and prose further down are the page's content, not a refusal; so is a heading the page's title names (an issue or a question whose subject begins with those words, "Please log in again #1411"). The rule matches English only. ladder_session_rejected names which: redirectedTo or signInPrompt. Import the site again to replace an expired login; a running server uses the new one at once.authed works on scrape and batch, not on crawl: a crawl follows every link, and a sign-out link would end your session in Chrome too, since the saved cookies are that session. Send the pages as a batch.localStorage (a token its script reads) needs a tab of it open when you import: Chrome reads an origin's storage only through a page that shows it, so with none open the import saves the cookies alone and says localStorageRead: false. A tab whose storage Chrome does not give (it crashed, was discarded or closed meanwhile) is saved without it and named in localStorageUnread, with the request to Chrome that failed and Chrome's answer in localStorageUnreadReasons. Only tabs in the profile the cookies come from are read, never an Incognito window's, and only each tab's own origin, not a frame of another origin inside it; sessionStorage and IndexedDB are not saved. Chrome's remote debugging reaches the default profile only.For people who do not want the technical switches, one option chooses how pages are reached: "access": "standard", "enhanced" or "my-browser" on POST /v1/scrape and POST /v1/batches (standard or enhanced on POST /v1/crawl), the MCP scrape, batch_scrape and crawl tools, and --access on the command line.
standard: Octocrawl's own fetching and its local browser. No rung that costs a third party runs; the routing audit says which it dropped (ladder_channels_filtered, reason access standard).enhanced: also what the server's access grant of tier enhanced approves: the paid providers below, in any mode, mode standard included. The grant's run budget caps a batch or a crawl; a single scrape calls each provider the server names at most once. A server without such a grant refuses it by name (unsupported_parameter).my-browser: your own Chrome, the same as "lane": "my-browser" (below). A crawl does not take it.Without access, a request runs as the server is configured. A batch or crawl keeps its choice, so a resumed run makes the same one. The per-route options below stay for developers.
By default a server uses none of the capabilities ADR 0005 puts behind a grant: no provider browser, no challenge solving, no provider stealth. An operator who wants them starts the server with a grant, a JSON file the server checks at startup; a grant with any problem stops startup with every problem listed.
octocrawl serve --access-grant grant.json
{
"tier": "enhanced",
"capabilities": ["vendor_remote_browser", "vendor_captcha_solving"],
"budget": { "perRunUsd": 5, "perRequestUsd": 0.5 },
"tariffs": { "browserbase": { "perHourUsd": 0.12, "maxSessionMs": 120000, "minBilledMs": 60000, "billingIncrementMs": 60000 } },
"attestation": { "principal": "you@example.com", "at": "2026-10-05T12:00:00Z", "statement": "I accept the provider's terms and the cost of these routes." }
}
W2L_ACCESS_GRANT takes the same, as a file path or the JSON itself.standard or my_browser grant may name only compatible_transport and egress_sessions. A capability that can cost money needs a positive perRunUsd, and an enhanced one needs an attestation.W2L_VENDORS with its key, in mode research or authed) are built only when the grant names vendor_remote_browser; the provider's challenge solving and stealth follow vendor_captcha_solving and vendor_stealth.ladder_step with escalate). When the stronger rungs then fail without a page (a network error, a provider's own error, the deadline), the page it stepped past is the answer (ladder_evidence_kept). It stops at a timeout (a slow or dead site would hold the scrape for the browser's wait too), at a rate limit (429, whose Retry-After a batch or crawl honours), at an error status that is the page's own answer (404, 410, 5xx), at your saved login's rung (the rungs after it do not carry the login), and after an earlier rung found content, which stays the answer.tariffs names, per provider, the prices you accepted from its pricing page (perCallUsd, perHourUsd) and what bounds a call's time: a per-hour price needs maxSessionMs, the longest session one call may hold; minBilledMs is the shortest time the provider bills a session, and billingIncrementMs the step it bills in (a minute, rounded up). A call's ceiling is perCallUsd plus perHourUsd for its longest session, at least minBilledMs and at least the shortest timeout the provider takes (60 s for Browserbase, 15 s for Steel), rounded up to the step. Under a tariff each call opens a connection and a session of its own, never shared with another call, ends it at maxSessionMs and releases it; the session is also created with the provider's own timeout and its proxies off, so it ends on the provider's side if the release fails, and bills no bandwidth. A per-GB price is refused: the bytes a provider's browser receives across every target it opens cannot be counted from here, so a bandwidth cost has no ceiling. A provider with no tariff is not called (ladder_channel_skipped, no price ceiling).perRunUsd (and the task's own cap) has left, so concurrent pages cannot pass it, and it is settled at the price the provider reported or, when it reported none (Browserbase and Steel state no price per request), at its ceiling, never below the cost; a call that threw or that the deadline cut is charged at its ceiling. perRequestUsd caps one page's calls within the run, and a single scrape's (else perRunUsd does). A call whose ceiling does not fit is skipped and the trace says so (budget); a task whose cap is used up stops (budget_exceeded with cost). The answer keeps usage.externalCostUsd for the exact cost (null when unknown) and adds externalCostChargedUsd, what the ledger charged all the run's calls, with a spend_settled trace event per call. A provider called with no budget at all is reserved and settled in an uncapped ledger of the run's own, so its calls are on the record too. Each page's Evidence Record lists them in access.paidCalls, in order: the provider (provider), the rung, the ADR 0005 capabilities its session was created with (capabilities), the ceiling reserved (ceilingUsd), what the ledger charged (chargedUsd), the price the provider stated (reportedCostUsd, null when it stated none), and outcome and reason, what Octocrawl made of the page the call returned by its own checks (a block page, an empty or unverified read, an identity it did not send), never the provider's word that it succeeded, null when the call returned no page; answer marks the call whose page is the record's. access.grant names the grant they were made under by the SHA-256 of its text (shasum -a 256 grant.json gives the same), its tier and its attestation time, never the attestation's principal or statement. A page read again on another egress keeps the calls of the read it gave up, a page you read in your Chrome after a check keeps those of the run the check stopped, and a batch or crawl page whose run threw after a paid call keeps it (none of these is its answer); a page from the cache lists those of the fetch it reuses. Both are null when no provider was called. A model's cost for JSON extraction is reported apart (modelUsage) and not counted. A grant that names scope.hosts is refused, since nothing limits the routes to those hosts yet.W2L_BROWSER_ENGINE=patchright runs the public browser rung on Patchright, a maintained Playwright fork, when the grant names enhanced_browser; a hosted server refuses it, and a saved login's rung and a managed session keep stock Playwright. Patchright is not installed with Octocrawl: install the two together (npm install octocrawl patchright, or npx -p octocrawl -p patchright octocrawl ...), then npx patchright install chromium. An executeJavascript step still runs in the page's own JavaScript world on it. A page fetched on it says so in its trace (browser_engine). It stays an experiment: on 2026-10-06, over one exit, it reached no more pages than stock Playwright and lost none (record), so Octocrawl does not turn it on by default.W2L_COMPAT_HOSTS=example.com,shop.example sends those hosts' pages, and their subdomains', over a browser-compatible HTTP transport (impit) when the grant names compatible_transport: the http rung becomes http_compat, in standard mode on a local server; a hosted server refuses it. The request carries Chrome's own headers and TLS handshake, so the page records that Chrome identity (identity_sent) and the transport (transport). A request with custom headers or mobile keeps the http rung, since the transport cannot send them without changing Chrome's header set, and so does a map's start page, read under the identity the map reports. impit decodes compressed bodies itself, so such a page reports its wire size as unknown. Without W2L_COMPAT_HOSTS, a grant that names compatible_transport uses it for the hosts its acceptance showed it helps: five hosts (research/access/benefit-hosts.v1.json) whose blocked task it verified in both G1 windows with no regression (research/access/runs/2026-10-07-g1-acceptance-5afe577.md). It is not for every host: across the whole task set it also turned refusals into answers whose data was wrong. W2L_COMPAT_HOSTS=none turns it off; naming hosts replaces the list.egress_sessions gives each batch and crawl (not a single scrape, nor mode authed, which has your saved login) a cookie session: the cookies its pages set are sent again to their site on its later pages, by the HTTP rung, the compatible transport and the browser alike, so a page whose check the browser cleared lets the task's next page of that site go over HTTP. Cookies are matched to domain and path as a browser does, kept in the task's directory (cookie-session.<route>.json, one per egress route, readable by you alone) so a batch or crawl resumed after a restart goes on with them, deleted when the task ends, and never recorded: a page's trace names the session's random id and counts (session_cookies). A page read with a session is not cached.W2L_EGRESS_PROXIES=http://user:pass@proxy-a:8080,http://proxy-b:3128 (with egress_sessions; a hosted server refuses it) sends every fetch through your own proxies. A batch or crawl keeps one for its run, and its cookie session belongs to that one; it moves to the next healthy proxy only when its proxy itself fails: after a page got no HTTP answer from any rung, Octocrawl asks the proxy for a tunnel to a name that does not exist, and moves on only if the proxy does not answer or refuses its credentials (407), at most twice a run, with a new cookie session, and the page is read again there (egress_switched in its trace). It never moves because of what a site did: a block, a challenge, a 429 or a connection the site reset stays on its proxy, since moving to another address to get past one is identity rotation, which Octocrawl does not do. A proxy that failed is set aside for 10 minutes; a scrape takes the next healthy one. A task resumed on another route (the pool or the proxy changed) starts a new cookie session; mode authed never uses the pool. Records name a proxy by host:port (proxy), never its credentials. With W2L_EGRESS_ECHO_URL set to a service that answers with the caller's address (https://ipinfo.io/json, https://api.ipify.org, https://httpbin.org/ip), each proxy is asked it through itself once, again after it failed or after 10 minutes, and every page read through that proxy records where it left from as access.egress.exit ({ ip, country, observedAt }, the country when the service gives one); without it, or when the echo did not answer, exit is null. The echo request goes to the service you name, through your proxy, on a connection of its own: a gateway that gives each connection or session a new exit may have sent the page from another address, so exit is where that proxy was seen to leave from, not proof of the page's own address. A page from the cache keeps the exit its original read recorded, and a page never waits for the echo past its timeout (the exit is then null).vendor_unlock_html and third_party_captcha_solver can be granted, but the routes that use them are not built yet (ROADMAP PA items 4 and 6).octocrawl commands read W2L_ACCESS_GRANT too, as their engine runs in their own process.By default Octocrawl does not solve captchas or challenges and does not disguise itself (see the access grant above for what a grant changes). When a page of a batch stops at one (blocked with captcha, cloudflare_challenge, bot_detected_generic or login_wall), Octocrawl on your own machine can hand it to you in the Chrome you already use: octocrawl batch <urls> --handoff, POST /v1/batches/:id/handoff on a local server (octocrawl serve on loopback), or the local MCP tool hand_off_batch. Remote debugging must be on, as for octocrawl login import, and Chrome asks "Allow remote debugging?" once per handoff. While remote debugging is on, every page Chrome opens sees navigator.webdriver as true, whether Octocrawl is connected or not (seen 2026-10-07 on Chrome 153 with the chrome://inspect switch, and on Chrome 154 started with --remote-debugging-port): a site that looks for it, as bot checks may, takes your Chrome for an automated one, and Chrome shows "Chrome is being controlled by automated test software". Octocrawl does not hide it, since it never changes your Chrome. Turn remote debugging off at chrome://inspect/#remote-debugging when you are done.
For one page, ask for it in the request: octocrawl scrape <url> --handoff, or "handoff": true (or { "waitMs": 60000 }) on POST /v1/scrape, the SDK's scrape and the MCP scrape tool. When Octocrawl's own fetch stops at such a check, the page opens in your Chrome in the same way, and the scrape answers with the page you get through to. That answer is recorded as below, and its routing audit is the stopped run's. A page that is not read answers as stopped, with a handoff_not_through warning that says why. The scrape's timeout bounds Octocrawl's fetch, not your time; the SDK waits for the answer as long as the handoff takes. Without the option, a stopped page on a server that offers the handoff carries handoff: { reason, liveViewUrl: null, rationale }, saying how to ask for it. A server that does not offer the handoff refuses the option (unsupported_parameter), as it does beside actions or a screenshot.
While the batch waits. It runs to the end as usual. The batch's status counts the stopped items in waitingForPerson. Each such item carries handoff: { reason, liveViewUrl: null, rationale }, where reason is captcha_required, bot_gate or login_required. A rate limit or a region block is not handed over.
What happens when you hand off. Each stopped page opens in a new tab of your Chrome, one at a time. You get through the check there, as you would on your own. Octocrawl then reads the page and closes the tab. A page counts as through when, on three reads a second apart, all of these hold:
The page's address is the one Chrome shows, not what the page's script says. A page that reloads itself (a challenge that runs a script, then reloads) is waited for, not taken for a closed tab. A way through that ends elsewhere on the site (a sign-in that lands on the home page) is followed by Octocrawl taking the tab back to the page asked for, twice at most; a tab still elsewhere after that is not read. The page asked for is the URL itself, or where the URL leads when Octocrawl takes the tab back to it, or that page with its address rewritten by its own script once it came, with no new document and before you clicked or typed on it (Indeed drops its paging token and names the job it shows): the same path, the same value for every parameter both addresses name, and no number dropped (a start=10 or page=2 that disappears may mean the site fell back to its first page). An address your own click or key moved in place (Next, a sort) is not the page asked for, nor is another page of a list; the tab is taken back.
Octocrawl reads a page in your Chrome only after you acted in its tab: a click or a key press that Chrome itself counts as a user's (the document's user activation, read in a world of Octocrawl's own that the page's script cannot reach, or a navigation Chrome marks as made with a user gesture). Nothing the page does by itself counts: not a reload, a redirect, a check that passes on its own, or a script filling a field. A page that shows no check in your Chrome (you are already signed in there, say) is read only once you click on it; octocrawl batch --handoff tells you so, and until you do the item keeps its stopped result. A click counts only in the tab Octocrawl opened: when that tab stays out of sight for 3 s (another tab or window in front of it), octocrawl scrape --handoff and octocrawl batch --handoff tell you to switch to it. So one Allow never lets a caller read the sites you are signed into without you; the my-browser lane below reads without a click only on the sites you allowed in Octocrawl's own page in your Chrome. To read pages with your login and no handoff, import it for the site (octocrawl login import) and run the batch in mode authed.
What replaces the stopped result. Only a read that is the page (success or partial) replaces the item's stopped result, in the formats the batch asked for, under the item's own id; a read that still shows a check, or is an error or empty, leaves the stopped result standing. The page is recorded as what it is: lane browser_local_authed, mode authed, compliance: null (Octocrawl sent nothing, so it signs nothing), usage.requestCount: 0, the User-Agent unknown (identity_unobserved: your browser sent the request), the document's status and Content-Type as your browser received them, and the robots.txt decision of the stopped fetch. The trace records the stopped result in handoff_from and the read in user_browser_read, with how you acted (act: user_activation or gesture_navigation) and the check Octocrawl saw (sawGate); the stopped run's routing audit is dropped with it.
When a page is not read. You may not get through in time (waitMs, 10 s to 30 min, default 10 min per page; a page left on another site, or one you did not click on, is given up when that time ends), close the tab or quit Chrome; or the caller may go away (the request's connection closes, Ctrl-C on the CLI), or Octocrawl may shut down; each ends the wait and closes the tab. A page Chrome refuses to open or answer for is not read, and the others are still handed over. That item keeps its stopped result, and the answer says why: { id, handedOff, through, notThrough, items: [{ id, url, through, status, reason? }] }.
A list that stopped at a check. For a batch whose only step is paginate with an itemSelector (the items are what tells a page of the list from another page you open in that tab), the page the check was on opens in your Chrome; for a pager whose pages have no address of their own, that is the list's own page, and the pages read before the check are shown again on the way without counting against maxPages. You get through it there, then page on yourself by clicking Next: Octocrawl only reads that tab (a read-only script; it clicks nothing and sends nothing), and stops when Next has been gone, hidden or disabled on the last page it read for 5 s, at the step's maxPages, after 60 s without a new page, or at the handoff's waitMs (default 10 minutes in all); wait for the terminal's or the tool's word that the check's page was read before you click Next, and do not close the tab (that, or quitting Chrome, gives up the item, pages read included); if you got through but showed no page after the check's within that time, the item keeps its stopped result and a later handoff goes on from there. The pages Octocrawl's own browser read before the check and the pages you showed it are merged into one list, each page once (actions.scrapes[].by is user_browser for yours); the item becomes the whole list, actions.lists[].continued is { from, pages, by: "user_browser" }, and stoppedBy says how the reading ended (end, max, or deadline with a list_not_exhausted warning).
Where it is offered. Only on a server that runs on your machine and answers you alone: a server not on loopback, and the hosted service, answer 409. A batch that asked for page actions or a screenshot is not handed over either (409, no handoff on its items): those are Octocrawl's browser's to take, and a page read in your Chrome cannot give them. Cancelling while Chrome still asks "Allow remote debugging?" drops the connection, so an Allow clicked later attaches to nothing. The handoff runs on a finished batch, not while it runs. It sends no webhook event for an item it replaces: read the items again.
octocrawl serve (and npm run api) reads saved logins only when it listens on loopback, and then answers only requests addressed to 127.0.0.1, localhost or [::1] with no foreign Origin, so another machine or a web page using a rebound DNS name cannot read pages as you. Listening on another address, it reads none and says so when it starts. A hosted server never reads them.
my-browser)Remote debugging, which this lane needs, makes every page see navigator.webdriver as true while it is on (see the handoff above). The one exception is a batch whose only step is paginate with an itemSelector: an item whose list stopped at a check (actions.lists[].stoppedBy is challenge) is handed over as above; its other stopped items are not. A page that only needs your login, your address or a real browser is read; a site whose bot check looks at navigator.webdriver may refuse your Chrome as it refuses Octocrawl's own lanes. In the 2026-10-07 acceptance run studylib.net and imf.org were read this way, while crunchbase.com (a Cloudflare block page) and stackoverflow.com (a Cloudflare challenge that did not clear) were not; why those two refused is not isolated.
On a server running on your machine, a scrape can read its page in your own Chrome instead of fetching it: octocrawl scrape <url> --lane my-browser, or "lane": "my-browser" on POST /v1/scrape, the SDK's scrape and the MCP scrape tool. Octocrawl fetches nothing itself; your Chrome loads the page, signed in as you where you are.
octocrawl login import). Octocrawl then opens a page of its own in your Chrome that lists the site and the task, and waits for you to click Allow reading these sites there. Only that click, which Chrome counts as yours, allows the site: a script cannot. Closing that page, or clicking Revoke, stops it, and a page being read then ends as cancelled. Not allowing the site in time (10 minutes), closing the page or clicking Revoke first refuses the request (409), and the refusal says what the page last answered (not clicked, or allowed without a click Chrome counted). A batch waiting for you to allow its sites says so in its status (waitingForApproval: true), and the server's log shows each step of the approval (my_browser_approval), never a page's content.display: none or visibility: hidden, such as a help panel or a tab not chosen, is left out; whether it shows a check, and its raw HTML, are the whole document. What a page draws on a canvas, as a Lark sheet draws its cells, is not in it to read. A page that shows a check waits for you to get through it (handoff.waitMs, default 10 min). A page not read answers cancelled, blocked with the check it still showed, failed/connection_error when Chrome refuses a command (a tab it will not open), failed/redirect_limit when the site leads the tab to another page of it each time Octocrawl takes it back (twice), or failed/timeout, with a my_browser_not_read warning that says why.my_browser, never cached; compliance: null and usage.requestCount: 0 (Octocrawl sent nothing), no robots.txt decision (it fetched nothing), and the Evidence Record's access says route: user_browser, the browser that read it, and completion: user_browser, or handed_to_person when the page showed a check you got through.actions, a screenshot, lockdown or a mode other than standard, the request is refused by name (unsupported_parameter)."lane": "my-browser" on POST /v1/batches, the MCP batch_scrape tool or octocrawl batch <urls> --lane my-browser reads every page that way, one at a time. Octocrawl's page in your Chrome lists every site of the batch (host and port) and you allow them once for the run; a page on a site not among them is not read, nor any page after you revoke them. Not allowing them answers every page cancelled, and Chrome not reached answers failed/connection_error, each with the reason. A run resumed later (a restart, a pause) asks you again: a server restarted while such a batch ran opens Chrome's prompt as it starts, and pages left when you do not answer end cancelled. A URL appended while the batch runs, on a site not in that run's list, ends cancelled too, and is not read later. Cancelling the batch, or its time budget ending, closes the tab being read and the connection. A page here has no timeout of its own (your time is yours), only the batch's. A batch on this lane takes no webhook (pages read as you are not sent elsewhere) and no maxConcurrency above 1.Every Evidence Record's access.completion counts how a page was read: unattended (Octocrawl's own lanes), authorized_session (your saved login, mode authed), user_browser (your Chrome, on a site you allowed, without a step of yours) or handed_to_person (your Chrome, after you got through a check). It is null when no page was read.
For researchers, two guides walk through a real run: From a URL list to a CSV with evidence (the command line and the Python client, every evidence column, and why failed rows stay) and Citing web data in a paper (a methods section, a reference with its access date and hash, and personal data).
The ports the local services listen on, all on 127.0.0.1:
| Port | Service | Started by | Changed with |
|---|---|---|---|
| 8787 | The REST API, and the API the stdio MCP server (octocrawl-mcp) and the Python client call by default (the TypeScript SDK takes a baseUrl) | octocrawl serve, npm run api | --port, W2L_API_PORT; those clients W2L_API_URL |
| 8791 | The local MCP service (API, Monitor scheduler, delivery worker and MCP endpoint at /mcp) | npm run local:mcp, or the LaunchAgent from npm run local:mcp:install | W2L_LOCAL_MCP_PORT |
| 8788 | The HTTPS webhook receiver of the first-use walkthrough | npm run first-use:local | fixed |
| 8798 | The public site's local preview | npm run public:preview:local | W2L_PUBLIC_PREVIEW_PORT |
For MCP use there are three ways, from least to most setup; the configs for Cursor, OpenCode and Codex, and a first task, are on Connect MCP.
Hosted (scrape and map; keyless within a daily allowance over HTTP, a key for more pages and the browser lane; see docs/hosted-api.md):
claude mcp add --transport http octocrawl https://mcp.octocrawl.dev/mcp
curl -sS -X POST https://api.octocrawl.dev/v1/scrape -H 'content-type: application/json' -d '{"url":"https://example.com"}'
On your computer (everything: scrape, map, crawl, batch, the Amazon.sg product tool and the Monitor tools; no limit; nothing from this checkout is needed):
npx octocrawl serve # keep it running: the API on 127.0.0.1:8787
claude mcp add octocrawl -- npx -y @octocrawl/mcp # the published stdio server, a client of that API
Self-hosted for others: npx octocrawl serve --hosted --token <token> listens on all interfaces behind a bearer token, with private addresses, robots overrides, saved logins, handoff and non-HTTPS webhooks refused; point @octocrawl/mcp at it with --base-url and --token (or W2L_API_URL and W2L_API_TOKEN).
The checkout's managed local service is for the Monitor → HTTPS delivery flow: one background service runs the API,
Monitor scheduler, delivery worker and a Streamable HTTP MCP endpoint at http://127.0.0.1:8791/mcp. On macOS,
install it as a LaunchAgent and connect Codex to its loopback URL:
npm run local:mcp:install
codex mcp add w2l-local --url http://127.0.0.1:8791/mcp
npm run local:mcp:status
It restarts after a process crash and at login. No hosting or sign-in account is
needed for this local path. npm run local:mcp:uninstall removes the agent;
codex mcp remove w2l-local removes the client entry. On other systems, run
npm run local:mcp in one terminal. The state stays in .w2l/api by default.
See the MCP first-use walkthrough for the actual
Monitor and HTTPS delivery flow and secret setup. Keep this checkout while
the LaunchAgent points to it.
To receive signed events on the same Mac with a fixed HTTPS loopback URL,
run npm run local:receiver:install, then reinstall the MCP service with
W2L_LOCAL_DELIVERY_LOOPBACK=1 npm run local:mcp:install. The option only
permits loopback delivery and pins trust to the generated local certificate.
The receiver and its SQLite inbox run as a separate LaunchAgent; neither
service becomes reachable from another machine.
The legacy standalone REST API remains available for SDK and Firecrawl-shim clients:
npm run api
To connect a standalone stdio MCP process to that API, run:
npm run mcp
npm run api binds 127.0.0.1 and allows loopback/RFC1918 so fixture servers work. Hosted mode is explicit: npm run api -- --hosted --token $W2L_API_TOKEN. That binds 0.0.0.0, requires Authorization: Bearer, denies private/metadata IPs, and limits a crawl to 100 pages: an omitted or null maxPages takes 100, and a larger one is refused with invalid_request. It obeys robots.txt for every URL and offers no way past it: robotsOverride, robotsOverrides and ignoreRobotsTxt are refused with unsupported_parameter (see below).
A server started with tokens, hosted or local, accepts any one of them: repeat --token, or set W2L_API_TOKEN and the comma-separated W2L_API_TOKENS. Tokens on the command line replace those in the environment. A --token without a value (the last argument, followed by another flag, or blank) stops the server at startup, and the error never repeats a token. Give each client its own token; restarting the server without a token revokes it. Tokens are compared as fixed-length SHA-256 digests in constant time, and a missing or unknown token gets HTTP 401 with { "error": "unauthorized", "code": "unauthorized" }. The SDK sends its token option, or W2L_API_TOKEN from the environment when none is passed; token: '' sends none.
An operator can cap how many requests that start work each caller may make: W2L_RATE_LIMIT_PER_MINUTE=<n> or --rate-limit-per-minute <n> (an integer from 1 to 100,000; unset or empty means no limit, anything else stops startup with W2L_RATE_LIMIT_PER_MINUTE must be an integer between 1 and 100000) counts POST /v1/scrape, /v1/crawl, /v1/batches, /v1/map, /fc/v1/scrape, /fc/v1/crawl and /fc/v1/map in a sliding 60-second window per bearer token (for the one local caller when the server takes no token); status reads are free. Over the limit the answer is HTTP 429 with a Retry-After header (whole seconds, at least 1) and { "error": "rate limit exceeded: <n> requests per minute", "code": "rate_limited", "retryAfterSeconds": <s>, "agentHints": ["wait <s> s before the next request"] }; /fc answers { success: false, error, code: "rate_limited", agent_hints } with the same header. rate_limited is not one of the request-error codes: the request was well formed, the caller's budget was spent. The SDK throws W2LError with status 429, code rate_limited, retryAfterMs read from the header (delta-seconds or an HTTP date) and agentHints, and retries nothing, as Firecrawl's SDKs do not; MCP tool calls fail with rate limited: retry after <s> s (rate_limited). The window is in memory and per process, so a restart resets it, and it keys on the token's digest, not the client address, so rotated tokens have separate budgets. Nothing about outbound politeness changes: the per-origin gate and the Retry-After cooldowns toward sites stay as they are.
Behind a proxy, local mode (npm run api, the local MCP service, npm run scrape/crawl) sends its outbound requests, including robots.txt and the local browser, through HTTPS_PROXY for https: URLs and HTTP_PROXY for http: URLs (lower-case names too), with curl's rules: NO_PROXY hosts and their subdomains, host:port, IP and CIDR entries go direct, * disables the proxy, and loopback is always direct. The proxy must be http:// or https://, and both variables must name the same one. The proxy resolves the names it fetches, so for proxied requests Octocrawl trusts it for resolution and checks only the URL itself (scheme, credentials, IP literals, metadata names); direct requests are still resolved, validated and pinned. Results name the proxy's host:port in evidence.envProxy and an egress_proxy trace event, never its credentials. W2L_PROXY=off ignores the variables; hosted mode never uses them. Without the environment proxy, the local browser connects directly like the HTTP lane: it never falls back to the operating system's proxy settings, a route no result would record. The macOS LaunchAgent does not inherit your shell, so put these variables in .w2l/local-mcp.env. Octocrawl verifies certificates by default, through the proxy too (the proxy tunnels TLS end to end); skipTlsVerification turns it off for one local request, is recorded, and is refused in hosted mode (see the scrape options below).
robots.txt is read for every URL Octocrawl fetches, and its verdict is recorded; what it decides depends on who chose the URL (decided 2026-10-05). robots.txt addresses crawlers that discover links, so on a local server a URL the request names (a scrape, a batch entry, the CLI's URL list, MCP scrape and batch_scrape, /fc/v1/scrape) is fetched whatever robots.txt says, as a browser visit would be: the result keeps the verdict (robotsDecision.decision: "disallowed", userOverride: true, overrideBasis: "user_named_url"), a robots_overridden warning and the trace events below. The links a crawl or map discovers obey robots.txt, unless the crawl or map was started with ignoreRobotsTxt on a local server (overrideBasis: "ignore_robots_txt"; a crawl then also reads the sitemap files robots.txt disallows, and a map returns the URLs it disallows, each link's robots saying disallowed or unreachable). A Monitor's scheduled re-reads obey it too. A hosted server obeys robots.txt for every URL, since it fetches from the operator's addresses. A site owner can address Octocrawl itself: robots.txt User-agent lines are matched against the request's User-Agent with the product token Octocrawl added, in every mode, so a group for Octocrawl (or research mode's w2l-research) governs it whatever header was sent. Such a rule is the owner's targeted opt-out: a named URL and ignoreRobotsTxt do not set it aside, and only a robotsOverride with your recorded reason does, on a local server. Pacing is unchanged by any of this: a batch and a crawl space a host's pages by its Crawl-delay, whether robots.txt allows the page or not, and a 429 cools the host down for every request. A 4xx robots.txt means no restrictions. A robots.txt that cannot be fetched (a 5xx, a n
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
npx -y @octocrawl/mcpMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-77777r7-octocrawl": {
"command": "npx",
"args": [
"-y",
"@octocrawl/mcp"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceio.github.77777R7/octocrawl works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.