Skip to content

Endpoints

Crawl

Whole sites, asynchronously. Stream pages live, poll with cursors or get webhooks.

POST/v1/crawl

ReferenceTry it ▸

Crawl starts an asynchronous job and returns immediately with its id. Discovery is sitemap-first, then follows links up to maxDepth. Every page is scraped with your scrapeOptions, and content that repeats across pages (navs, footers, cookie banners) is removed crawl-wide by default, which no single-page scraper can do. 1 credit per page scraped successfully (more if a page needs premium); failed pages are free.

Start a crawl#

curl -X POST "https://api.vermin.dev/v1/crawl" \
  -H "Authorization: Bearer $VERMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://docs.example.com",
    "limit": 500,
    "includePaths": ["/guides/**"],
    "scrapeOptions": { "formats": ["markdown"] }
  }'
202 Accepted
{ "id": "crawl_01J9V4A7", "status": "queued", "url": "https://docs.example.com", "createdAt": "2026-09-28T09:40:00Z" }
CrawlRequest
  • urlstring (uri)required
  • limitintegerdefault 1001–100000
  • maxDepthinteger0–20
    Link hops from the start URL (sitemap URLs count as depth 1). 0 = only the start URL.
  • includePathsstring[]
    Path globs matched against path + query: * within a segment, ** across segments, a trailing $ anchors the end; patterns starting with / match from the start of the path.
  • excludePathsstring[]
    Path globs (same syntax as includePaths) never crawled.
  • allowSubdomainsbooleandefault false
  • sitemap"include" | "skip" | "only"default "include"
  • removeBoilerplatebooleandefault true
    Strip content repeated across pages (nav, footers)
  • ignoreRobotsbooleandefault false
  • scrapeOptionsScrapeOptions
  • webhookobject
    Deliver crawl events to this URL as signed POSTs (see the crawlEvent webhook). Deliveries are signed with your organisation's webhook signing secret (Dashboard → Webhooks); endpoints registered in the dashboard also receive crawl.* events, signed with their own secret.
    Show 3 fields
    • urlstring (uri)
    • events("started" | "page" | "completed" | "failed" | "cancelled")[]
      Default: all events.
    • metadataobject
      Echoed back in every delivery.

Path globs#

includePaths and excludePaths match against the path plus query string: * matches within a segment, ** across segments, a trailing $ anchors the end, and patterns starting with / match from the start of the path. /blog/** keeps the blog; *.pdf$ skips PDFs.

Follow progress#

Three ways, pick whichever suits your architecture.

1. Stream it (server-sent events)#

GET/v1/crawl/{id}/events

ReferenceTry it ▸

The stream replays everything so far, then follows the crawl live and closes after done. Events: page (a Document), status (counts) and done (final status). Every event has an id; reconnect with Last-Event-ID and you'll pick up exactly where you left off. The SDKs do this for you.

curl -N "https://api.vermin.dev/v1/crawl/crawl_01J9V4A7/events" \
  -H "Authorization: Bearer $VERMIN_API_KEY" \
  -H "Accept: text/event-stream"
text/event-stream
id: 1
event: status
data: {"id":"crawl_01J9V4A7","status":"running","total":212,"completed":0,"failed":0}

id: 2
event: page
data: {"url":"https://docs.example.com/guides/intro","markdown":"# Intro\n...","tier":"fetch"}

id: 214
event: done
data: {"id":"crawl_01J9V4A7","status":"completed","total":212,"completed":209,"failed":3,"credits_used":209}

2. Poll with cursors#

GET/v1/crawl/{id}

ReferenceTry it ▸

Returns counts plus up to limit pages. Keep calling with cursor: next until next is null.

curl "https://api.vermin.dev/v1/crawl/crawl_01J9V4A7?limit=50" \
  -H "Authorization: Bearer $VERMIN_API_KEY"
CrawlStatus
  • idstringrequired
  • status"queued" | "running" | "completed" | "failed" | "cancelled"required
  • urlstring
  • createdAtstring (date-time)
  • totalinteger
  • completedinteger
  • failedinteger
  • credits_usedinteger
  • dataDocument[]
  • nextstring | null
    Cursor for the next page of results; null once the crawl is finished and every page has been returned.
  • errorobject
    Why the crawl failed or stopped early (e.g. spend_cap_reached, insufficient_credits, blocked_by_site). Pages crawled before the stop are still returned.
    Show 2 fields
    • codestringrequired
    • messagestringrequired

3. Webhooks#

Add webhook: { url, events, metadata } to the request and we'll POST signed events to you as the crawl runs. See Webhooks & streaming for payloads and signature verification.

Cancel#

DELETE/v1/crawl/{id}

ReferenceTry it ▸

Stops scheduling new pages. Pages already scraped are kept (and billed); nothing else is.

curl -X DELETE "https://api.vermin.dev/v1/crawl/crawl_01J9V4A7" \
  -H "Authorization: Bearer $VERMIN_API_KEY"