Skip to content

Endpoints

Scrape

One URL in, clean markdown, HTML, screenshots, chunks, metadata or brand out.

POST/v1/scrape

ReferenceTry it ▸

Scrape fetches a single URL and returns the formats you ask for. With the default tier: "auto" it tries a plain fetch, runs a quality check, and only escalates to a headless browser when the page actually needs one. You pay for the tier that worked: 1 credit (rendering and screenshots included), +1 when you use browser actions, 15 total on the premium proxy tier.

Example#

curl -X POST "https://api.vermin.dev/v1/scrape" \
  -H "Authorization: Bearer $VERMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/blog/post",
    "formats": ["markdown", "chunks", "screenshot"],
    "chunkSize": 800
  }'

Request#

ScrapeRequest
  • urlstring (uri)required
  • formatsFormat[]default ["markdown","metadata"]
    Basic metadata (title, description, language, canonical, status) is always returned; metadata adds parsed JSON-LD.
  • onlyMainContentbooleandefault true
  • includeTagsstring[]
  • excludeTagsstring[]
  • tier"auto" | "fetch" | "render" | "premium"default "auto"
    auto escalates fetch → render → (premium if allowed)
  • allowPremiumbooleandefault false
  • actionsAction[]max 20 items
  • waitForinteger≤ 30000
  • timeoutintegerdefault 300001000–60000
  • maxAgeintegerdefault 172800000
    Serve a cached copy younger than this (ms). Default 2 days; 0 forces a fresh scrape. Cache hits are served from the nearest data center.
  • headersRecord<string, string>
  • locationobject
    Show 2 fields
    • countrystring
    • languagesstring[]
  • chunkSizeintegerdefault 1500200–8000

Formats#

FormatYou get
markdownMain content as GitHub-flavoured markdown (default).
htmlCleaned HTML of the main content.
rawHtmlThe page's HTML exactly as served or rendered.
textPlain text, no markup.
linksEvery absolute link on the page.
screenshot / screenshotFullSigned image URL of the viewport or the full page (expires in 24h). Forces the render tier, at no extra cost.
chunksMarkdown split on headings into ~chunkSize-token pieces, each with its headingPath. Ready for your vector store.
metadataAdds the page's structured data (JSON-LD) to metadata. Title, description, language, canonical URL, status code, author, dates and OG image come back on every scrape, whatever the formats.
brandThe page's brand identity, extracted in the same pass.

Actions#

Need to click a cookie banner, scroll a feed or type into a search box first? Pass up to 20 actions; they run in order in a real browser before the page is captured.

actions
{
  "url": "https://example.com/feed",
  "actions": [
    { "type": "click", "selector": "#accept-cookies" },
    { "type": "scroll" },
    { "type": "wait", "ms": 1500 }
  ]
}

Caching#

Public pages are cached for two days by default and a cache hit costs 1 credit, served from the nearest data center in milliseconds (cached: true). Set maxAge (ms) to decide how fresh is fresh enough, or maxAge: 0 to always scrape live. Requests with custom headers, actions or screenshot formats are never cached or served from cache, and neither are pages the site marks no-store or private.

Response#

200 OK
{
  "success": true,
  "data": {
    "url": "https://example.com/blog/post",
    "finalUrl": "https://example.com/blog/post",
    "markdown": "# Shipping at the edge\n\n...",
    "chunks": [
      { "index": 0, "text": "# Shipping at the edge ...", "headingPath": ["Shipping at the edge"], "tokens": 781 }
    ],
    "screenshot": "https://api.vermin.dev/artifacts/screenshots/7f3c….png?sig=…",
    "metadata": { "title": "Shipping at the edge", "language": "en", "statusCode": 200 },
    "tier": "render",
    "cached": false,
    "quality": {
      "score": 0.93,
      "issues": ["js_required"],
      "escalations": ["fetch→render: js_required"],
      "reason": "The fetched HTML was an empty app shell; the rendered page had the article."
    }
  },
  "credits_used": 1,
  "latency_ms": 1240,
  "request_id": "req_01J9V3P0ZC"
}
Document
  • urlstringrequired
  • finalUrlstring
  • markdownstring
  • htmlstring
  • rawHtmlstring
  • textstring
  • linksstring[]
  • screenshotstring
    Signed URL (expires in 24h)
  • chunksChunk[]
  • metadataMetadata
  • brandBrand
  • tier"fetch" | "render" | "premium"
  • cachedboolean
  • qualityQuality
    Why content may be thin, and what the pipeline did about it.

The quality report#

quality is Vermin's answer to "why is my markdown empty?":

  • score: 0–1 confidence that the content is complete and is the page's real main content.
  • issues: what we noticed. thin_content, js_required, bot_wall, cookie_wall, login_wall, soft_404, truncated, non_html.
  • escalations: what the pipeline did about it, e.g. fetch→render: js_required.
  • reason: the same story in words.

Errors#

Common ones: blocked_url (private/local address or robots.txt), blocked_by_site, timeout, capacity_exceeded, and not_configured if you ask for a tier this deployment doesn't have. All free. See Errors.