Endpoints
Scrape
One URL in, clean markdown, HTML, screenshots, chunks, metadata or brand out.
Scrape fetches a single URL and returns the formats you ask for. With the default tier: "auto" it tries a plain fetch, runs a quality check, and only escalates to a headless browser when the page actually needs one. You pay for the tier that worked: 1 credit (rendering and screenshots included), +1 when you use browser actions, 15 total on the premium proxy tier.
Example#
curl -X POST "https://api.vermin.dev/v1/scrape" \
-H "Authorization: Bearer $VERMIN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/blog/post",
"formats": ["markdown", "chunks", "screenshot"],
"chunkSize": 800
}'Request#
urlstring (uri)required- Basic metadata (title, description, language, canonical, status) is always returned;
metadataadds parsed JSON-LD. onlyMainContentbooleandefault trueincludeTagsstring[]excludeTagsstring[]tier"auto" | "fetch" | "render" | "premium"default "auto"auto escalates fetch → render → (premium if allowed)allowPremiumbooleandefault falsewaitForinteger≤ 30000timeoutintegerdefault 300001000–60000maxAgeintegerdefault 172800000Serve a cached copy younger than this (ms). Default 2 days; 0 forces a fresh scrape. Cache hits are served from the nearest data center.headersRecord<string, string>locationobjectShowHide 2 fields
countrystringlanguagesstring[]
chunkSizeintegerdefault 1500200–8000
Formats#
| Format | You get |
|---|---|
markdown | Main content as GitHub-flavoured markdown (default). |
html | Cleaned HTML of the main content. |
rawHtml | The page's HTML exactly as served or rendered. |
text | Plain text, no markup. |
links | Every absolute link on the page. |
screenshot / screenshotFull | Signed image URL of the viewport or the full page (expires in 24h). Forces the render tier, at no extra cost. |
chunks | Markdown split on headings into ~chunkSize-token pieces, each with its headingPath. Ready for your vector store. |
metadata | Adds the page's structured data (JSON-LD) to metadata. Title, description, language, canonical URL, status code, author, dates and OG image come back on every scrape, whatever the formats. |
brand | The page's brand identity, extracted in the same pass. |
Actions#
Need to click a cookie banner, scroll a feed or type into a search box first? Pass up to 20 actions; they run in order in a real browser before the page is captured.
{
"url": "https://example.com/feed",
"actions": [
{ "type": "click", "selector": "#accept-cookies" },
{ "type": "scroll" },
{ "type": "wait", "ms": 1500 }
]
}Caching#
Public pages are cached for two days by default and a cache hit costs 1 credit, served from the nearest data center in milliseconds (cached: true). Set maxAge (ms) to decide how fresh is fresh enough, or maxAge: 0 to always scrape live. Requests with custom headers, actions or screenshot formats are never cached or served from cache, and neither are pages the site marks no-store or private.
Response#
{
"success": true,
"data": {
"url": "https://example.com/blog/post",
"finalUrl": "https://example.com/blog/post",
"markdown": "# Shipping at the edge\n\n...",
"chunks": [
{ "index": 0, "text": "# Shipping at the edge ...", "headingPath": ["Shipping at the edge"], "tokens": 781 }
],
"screenshot": "https://api.vermin.dev/artifacts/screenshots/7f3c….png?sig=…",
"metadata": { "title": "Shipping at the edge", "language": "en", "statusCode": 200 },
"tier": "render",
"cached": false,
"quality": {
"score": 0.93,
"issues": ["js_required"],
"escalations": ["fetch→render: js_required"],
"reason": "The fetched HTML was an empty app shell; the rendered page had the article."
}
},
"credits_used": 1,
"latency_ms": 1240,
"request_id": "req_01J9V3P0ZC"
}The quality report#
quality is Vermin's answer to "why is my markdown empty?":
score: 0–1 confidence that the content is complete and is the page's real main content.issues: what we noticed.thin_content,js_required,bot_wall,cookie_wall,login_wall,soft_404,truncated,non_html.escalations: what the pipeline did about it, e.g.fetch→render: js_required.reason: the same story in words.
Errors#
Common ones: blocked_url (private/local address or robots.txt), blocked_by_site, timeout, capacity_exceeded, and not_configured if you ask for a tier this deployment doesn't have. All free. See Errors.