Endpoints
Crawl
Whole sites, asynchronously. Stream pages live, poll with cursors or get webhooks.
Crawl starts an asynchronous job and returns immediately with its id. Discovery is sitemap-first, then follows links up to maxDepth. Every page is scraped with your scrapeOptions, and content that repeats across pages (navs, footers, cookie banners) is removed crawl-wide by default, which no single-page scraper can do. 1 credit per page scraped successfully (more if a page needs premium); failed pages are free.
Start a crawl#
curl -X POST "https://api.vermin.dev/v1/crawl" \
-H "Authorization: Bearer $VERMIN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://docs.example.com",
"limit": 500,
"includePaths": ["/guides/**"],
"scrapeOptions": { "formats": ["markdown"] }
}'{ "id": "crawl_01J9V4A7", "status": "queued", "url": "https://docs.example.com", "createdAt": "2026-09-28T09:40:00Z" }urlstring (uri)requiredlimitintegerdefault 1001–100000maxDepthinteger0–20Link hops from the start URL (sitemap URLs count as depth 1). 0 = only the start URL.includePathsstring[]Path globs matched against path + query:*within a segment,**across segments, a trailing$anchors the end; patterns starting with/match from the start of the path.excludePathsstring[]Path globs (same syntax as includePaths) never crawled.allowSubdomainsbooleandefault falsesitemap"include" | "skip" | "only"default "include"removeBoilerplatebooleandefault trueStrip content repeated across pages (nav, footers)ignoreRobotsbooleandefault falsescrapeOptionsScrapeOptionswebhookobjectDeliver crawl events to this URL as signed POSTs (see thecrawlEventwebhook). Deliveries are signed with your organisation's webhook signing secret (Dashboard → Webhooks); endpoints registered in the dashboard also receivecrawl.*events, signed with their own secret.ShowHide 3 fields
urlstring (uri)events("started" | "page" | "completed" | "failed" | "cancelled")[]Default: all events.metadataobjectEchoed back in every delivery.
Path globs#
includePaths and excludePaths match against the path plus query string: * matches within a segment, ** across segments, a trailing $ anchors the end, and patterns starting with / match from the start of the path. /blog/** keeps the blog; *.pdf$ skips PDFs.
Follow progress#
Three ways, pick whichever suits your architecture.
1. Stream it (server-sent events)#
The stream replays everything so far, then follows the crawl live and closes after done. Events: page (a Document), status (counts) and done (final status). Every event has an id; reconnect with Last-Event-ID and you'll pick up exactly where you left off. The SDKs do this for you.
curl -N "https://api.vermin.dev/v1/crawl/crawl_01J9V4A7/events" \
-H "Authorization: Bearer $VERMIN_API_KEY" \
-H "Accept: text/event-stream"id: 1
event: status
data: {"id":"crawl_01J9V4A7","status":"running","total":212,"completed":0,"failed":0}
id: 2
event: page
data: {"url":"https://docs.example.com/guides/intro","markdown":"# Intro\n...","tier":"fetch"}
id: 214
event: done
data: {"id":"crawl_01J9V4A7","status":"completed","total":212,"completed":209,"failed":3,"credits_used":209}2. Poll with cursors#
Returns counts plus up to limit pages. Keep calling with cursor: next until next is null.
curl "https://api.vermin.dev/v1/crawl/crawl_01J9V4A7?limit=50" \
-H "Authorization: Bearer $VERMIN_API_KEY"idstringrequiredstatus"queued" | "running" | "completed" | "failed" | "cancelled"requiredurlstringcreatedAtstring (date-time)totalintegercompletedintegerfailedintegercredits_usedintegerdataDocument[]nextstring | nullCursor for the next page of results; null once the crawl is finished and every page has been returned.errorobjectWhy the crawl failed or stopped early (e.g. spend_cap_reached, insufficient_credits, blocked_by_site). Pages crawled before the stop are still returned.ShowHide 2 fields
codestringrequiredmessagestringrequired
3. Webhooks#
Add webhook: { url, events, metadata } to the request and we'll POST signed events to you as the crawl runs. See Webhooks & streaming for payloads and signature verification.
Cancel#
Stops scheduling new pages. Pages already scraped are kept (and billed); nothing else is.
curl -X DELETE "https://api.vermin.dev/v1/crawl/crawl_01J9V4A7" \
-H "Authorization: Bearer $VERMIN_API_KEY"