Skip to content

Product

Everything you need to read the web programmatically.

Four endpoints on one retrieval engine. One key, one credit balance. Ask for the data; Vermin works out how to get it.

1 credit / page

Scrape

Fetch one URL and get clean markdown, HTML, links, chunks and metadata back.

One page, picked clean.

Scrape reference →
request
curl -X POST "https://api.vermin.dev/v1/scrape" \
  -H "Authorization: Bearer $VERMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://shop.example/p/42", "formats": ["markdown", "chunks"] }'
response (trimmed)
{
  "markdown": "# Rat-proof storage bins\n\nMade from 1.2 mm steel…",
  "chunks": [{ "headingPath": ["Rat-proof storage bins"], "tokens": 318 }],
  "tier": "render",
  "credits_used": 1
}

1 credit / page

Crawl

Follow links across a whole site and collect every page, with boilerplate removed site-wide.

Breeds across the whole site.

Crawl reference →
request
curl -X POST "https://api.vermin.dev/v1/crawl" \
  -H "Authorization: Bearer $VERMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://docs.example.com",
    "limit": 500,
    "includePaths": ["/guides/**"]
  }'
response (trimmed)
{ "id": "crawl_01J9V4A7", "status": "queued" }

// stream: GET /v1/crawl/crawl_01J9V4A7/events
{ "url": ".../guides/auth", "markdown": "# Auth…", "tier": "fetch" }

3 credits / 10 fetches

Map

List every URL on a site from sitemaps, robots and the link graph, ranked by a query.

Knows every tunnel.

Map reference →
request
curl -X POST "https://api.vermin.dev/v1/map" \
  -H "Authorization: Bearer $VERMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://docs.example.com",
    "search": "authentication",
    "limit": 100
  }'
response (trimmed)
{
  "links": [
    { "url": ".../guides/auth", "title": "Authentication", "source": "sitemap" },
    { "url": ".../api/keys", "title": "API keys", "source": "link" }
  ]
}

5 credits · 1 cached

Brand

Turn a domain into its name, logo, icon, colour palette, fonts and socials.

Steals the logo. And the colours.

Brand reference →
request
curl "https://api.vermin.dev/v1/brand?domain=stripe.com" \
  -H "Authorization: Bearer $VERMIN_API_KEY"
response (trimmed)
{
  "name": "Stripe",
  "logo": { "format": "svg", "background": "light" },
  "colors": [{ "hex": "#635BFF", "role": "primary", "textColor": "#FFFFFF" }],
  "fonts": [{ "family": "Sohne", "role": "heading" }]
}

Escalation

Plain fetch first. A browser only when needed.

A quality check spots empty JavaScript shells, cookie walls, soft 404s and bot walls. Rendering costs nothing extra; browser actions (click, scroll, wait) add 1. Every response says which tier it used and why.

  1. 011 credit

    Fetch

    Plain HTTP, then a quality check. Fast, cheap, and enough for most of the web.

  2. 02same price

    Browser

    Empty JavaScript shell? A headless browser renders it and waits for the content.

  3. 030 credits

    Still blocked?

    Bot wall, timeout, 404. You get a clear error code and keep your credit.

Output

Formats your code can use.

markdown, html, text, links, metadata, screenshot and chunks with heading paths and token counts. PDFs, DOCX, JSON and feeds too. Crawls remove boilerplate across the whole site.

data.chunks[0]
{
  "index": 0,
  "headingPath": ["Rat-proof storage bins", "Sizes"],
  "text": "| Size | Litres | Price |\n|---|---|---|…",
  "tokens": 204
}

Manners

Gets in everywhere it's allowed.

Scrapes, crawls and maps respect robots.txt by default. It identifies itself as VerminBot, paces itself per host, and refuses private and internal addresses.