SDKs
Python SDK
The official Vermin client for Python. This page is the SDK's README, rendered.
Official Python SDK for Vermin, the web-data API for AI: scrape, crawl, map, search, extract and brand.
- Sync
Verminand asyncAsyncVerminclients with the same methods, built on httpx. - Typed pydantic v2 response models that follow the OpenAPI contract.
- Automatic retries with jittered backoff on 429, 5xx and network errors. They honour
retry_after_ms/Retry-After. - Auto-generated
Idempotency-Keyonscrapeandcrawl.start, reused across retries. credits_usedandrequest_idon every result, plus a typedVerminError.
Requires Python 3.9+.
pip install verminQuick start#
from vermin import Vermin
vermin = Vermin() # reads VERMIN_API_KEY (and optionally VERMIN_API_URL)
page = vermin.scrape("https://example.com", formats=["markdown", "metadata"])
print(page.data.markdown, page.data.metadata.title)
print(f"cost: {page.credits_used} credits (request {page.request_id})")Async:
import asyncio
from vermin import AsyncVermin
async def main():
async with AsyncVermin() as vermin:
page = await vermin.scrape("https://example.com")
print(page.data.markdown)
asyncio.run(main())Options#
Vermin(
api_key="vmn_live_…", # default: VERMIN_API_KEY
base_url="https://api.vermin.dev", # default: VERMIN_API_URL, then https://api.vermin.dev
timeout=60.0, # per-attempt HTTP timeout, seconds
max_retries=2, # retries on 429 / 5xx / network errors
headers={"X-Team": "growth"}, # extra headers on every request
http_client=httpx.Client(proxy=…), # bring your own httpx client
)Use the client as a context manager (with Vermin() as vermin:), or call close() / await aclose() when you're done.
Keyword options use snake_case (only_main_content=True). The SDK sends them as the API's camelCase. Nested scrape_options accept a dict with snake_case or camelCase keys, or a ScrapeOptions model.
API#
scrape(url, **options)#
res = vermin.scrape(
"https://news.ycombinator.com",
formats=["markdown", "links", "chunks"],
only_main_content=True,
tier="auto", # fetch → render → premium (if allow_premium)
chunk_size=1200,
actions=[{"type": "click", "selector": "#more"}, {"type": "wait", "ms": 1000}],
)
for chunk in res.data.chunks or []:
print(chunk.heading_path, chunk.tokens)
res.data.quality.issues # e.g. ["js_required"]crawl#
job = vermin.crawl.start("https://docs.example.com", limit=200, max_depth=3,
scrape_options={"formats": ["markdown"]})
# Poll until finished and collect every page (follows `next` cursors):
result = vermin.crawl.wait(job.id, poll_interval=2.0, timeout=600,
on_status=lambda s: print(s.completed, "/", s.total))
print(result.status, len(result.data), result.credits_used)
# …or stream pages as they arrive (server-sent events):
for event in vermin.crawl.stream(job.id):
if event.type == "page":
print(event.data.url)
elif event.type == "done":
print("finished:", event.data.status)
# Lower level:
status = vermin.crawl.get(job.id, limit=50)
page2 = vermin.crawl.get(job.id, cursor=status.next) if status.next else None
for doc in vermin.crawl.documents(job.id):
print(doc.url)
vermin.crawl.cancel(job.id)With AsyncVermin, use await vermin.crawl.wait(...), async for event in vermin.crawl.stream(...) and async for doc in vermin.crawl.documents(...).
crawl.wait returns the final status even if the crawl failed or was cancelled, so check result.status. It raises VerminError(code="timeout") if timeout seconds pass first. A failed crawl carries result.error (code, message). crawl.stream reconnects with Last-Event-ID if the connection drops before done (waiting the server's retry: delay, else backoff), up to max_reconnects consecutive times (default max_retries; the budget resets after each page/status event), then raises VerminError(code="connection_error"). Resume with crawl.stream(id, last_event_id="41").
map(url, **options)#
res = vermin.map("https://example.com", search="pricing", limit=500)
urls = [link.url for link in res.links]
vermin.map("https://example.com", include_paths=["/blog/*"], exclude_paths=["/blog/tag/*"]) # path globssearch(query, **options)#
res = vermin.search("best vector databases 2026", limit=5, time_range="month",
scrape_options={"formats": ["markdown"]}) # optional: scrape each result
for hit in res.results:
print(hit.position, hit.title, hit.document and len(hit.document.markdown or ""))
# hit.date and hit.source are set when the provider reports them (e.g. news results)extract(urls, schema=..., prompt=...)#
Pass a pydantic model class or a JSON Schema dict. A model is sent as its JSON Schema. The returned data is then parsed into an instance of the model.
from pydantic import BaseModel
class Product(BaseModel):
name: str
price: float
in_stock: bool
res = vermin.extract("https://shop.example.com/p/123", schema=Product, prompt="Extract the product")
res.data.price # float, and res.data is a Product
raw = vermin.extract(["https://example.com"], schema={"type": "object", "properties": {"title": {"type": "string"}}})
raw.data["title"]If the data doesn't match the model, extract raises VerminError(code="invalid_response"). Pass validate=False to get the raw dict instead.
brand(domain, max_age=None)#
res = vermin.brand("stripe.com") # max_age=0 forces a fresh extraction
brand = res.data
brand.logo.url # Vermin-hosted copy; brand.logo.source_url is where it was found
next(c.hex for c in brand.colors if c.role == "primary") # each colour has text_color + WCAG contrast
# Every image found, tagged with type ("logo" | "icon" | "wordmark") and background ("light" | "dark" | "any")
for_dark_ui = [l for l in brand.logos or [] if l.background == "dark" and l.type != "icon"]
res.cached, res.stale # stale: cached result older than max_age, served because a fresh lookup failedErrors#
Every failure raises VerminError:
from vermin import VerminError
try:
vermin.scrape("https://example.com")
except VerminError as err:
err.code # "rate_limited", "insufficient_credits", "blocked_by_site", … or "connection_error"
err.status # HTTP status (0 if no response)
err.request_id # quote this to support
err.retry_after_ms # server hint, when present
err.retryable # whether retrying could helpRetries apply to 429, 5xx (including capacity_exceeded), timeouts and connection errors. By default the SDK makes up to 3 attempts, with jittered exponential backoff from 0.5 s up to 8 s. When the server sends a retry_after_ms hint (or a Retry-After header, seconds or HTTP date), the SDK waits exactly that long. It never retries invalid_request, unauthorized, forbidden, not_found, insufficient_credits, spend_cap_reached, blocked_url, not_implemented, not_configured or extract_failed, whatever the status. Failed requests cost 0 credits.
Webhooks#
Crawl webhooks carry a Vermin-Signature: t=<unix seconds>,v1=<hex> header, where v1 is the HMAC-SHA256 of "<t>.<raw body>" keyed with your signing secret (Dashboard → Webhooks). Verify it against the raw request body before trusting the payload. The helper compares in constant time, accepts any of several v1= values (secret rotation) and rejects timestamps more than 300 s from now (tolerance; 0 disables the check).
import os
from vermin import VerminError, verify_webhook
def handle(body: bytes, headers) -> int:
try:
event = verify_webhook(body, headers.get("Vermin-Signature"), os.environ["VERMIN_WEBHOOK_SECRET"])
except VerminError: # code == "invalid_signature"
return 400
if event["type"] == "crawl.page":
print(event["data"]["url"])
return 204is_valid_webhook() returns a bool instead, and sign_webhook() builds a header for your tests.