Skip to content

SDKs

Python SDK

The official Vermin client for Python. This page is the SDK's README, rendered.

verminpip install vermin

Official Python SDK for Vermin, the web-data API for AI: scrape, crawl, map, search, extract and brand.

  • Sync Vermin and async AsyncVermin clients with the same methods, built on httpx.
  • Typed pydantic v2 response models that follow the OpenAPI contract.
  • Automatic retries with jittered backoff on 429, 5xx and network errors. They honour retry_after_ms / Retry-After.
  • Auto-generated Idempotency-Key on scrape and crawl.start, reused across retries.
  • credits_used and request_id on every result, plus a typed VerminError.

Requires Python 3.9+.

bash
pip install vermin

Quick start#

python
from vermin import Vermin

vermin = Vermin()  # reads VERMIN_API_KEY (and optionally VERMIN_API_URL)

page = vermin.scrape("https://example.com", formats=["markdown", "metadata"])
print(page.data.markdown, page.data.metadata.title)
print(f"cost: {page.credits_used} credits (request {page.request_id})")

Async:

python
import asyncio
from vermin import AsyncVermin

async def main():
    async with AsyncVermin() as vermin:
        page = await vermin.scrape("https://example.com")
        print(page.data.markdown)

asyncio.run(main())

Options#

python
Vermin(
    api_key="vmn_live_…",               # default: VERMIN_API_KEY
    base_url="https://api.vermin.dev",  # default: VERMIN_API_URL, then https://api.vermin.dev
    timeout=60.0,                       # per-attempt HTTP timeout, seconds
    max_retries=2,                      # retries on 429 / 5xx / network errors
    headers={"X-Team": "growth"},       # extra headers on every request
    http_client=httpx.Client(proxy=…),  # bring your own httpx client
)

Use the client as a context manager (with Vermin() as vermin:), or call close() / await aclose() when you're done.

Keyword options use snake_case (only_main_content=True). The SDK sends them as the API's camelCase. Nested scrape_options accept a dict with snake_case or camelCase keys, or a ScrapeOptions model.

API#

scrape(url, **options)#

python
res = vermin.scrape(
    "https://news.ycombinator.com",
    formats=["markdown", "links", "chunks"],
    only_main_content=True,
    tier="auto",                 # fetch → render → premium (if allow_premium)
    chunk_size=1200,
    actions=[{"type": "click", "selector": "#more"}, {"type": "wait", "ms": 1000}],
)
for chunk in res.data.chunks or []:
    print(chunk.heading_path, chunk.tokens)
res.data.quality.issues  # e.g. ["js_required"]

crawl#

python
job = vermin.crawl.start("https://docs.example.com", limit=200, max_depth=3,
                         scrape_options={"formats": ["markdown"]})

# Poll until finished and collect every page (follows `next` cursors):
result = vermin.crawl.wait(job.id, poll_interval=2.0, timeout=600,
                           on_status=lambda s: print(s.completed, "/", s.total))
print(result.status, len(result.data), result.credits_used)

# …or stream pages as they arrive (server-sent events):
for event in vermin.crawl.stream(job.id):
    if event.type == "page":
        print(event.data.url)
    elif event.type == "done":
        print("finished:", event.data.status)

# Lower level:
status = vermin.crawl.get(job.id, limit=50)
page2 = vermin.crawl.get(job.id, cursor=status.next) if status.next else None
for doc in vermin.crawl.documents(job.id):
    print(doc.url)
vermin.crawl.cancel(job.id)

With AsyncVermin, use await vermin.crawl.wait(...), async for event in vermin.crawl.stream(...) and async for doc in vermin.crawl.documents(...).

crawl.wait returns the final status even if the crawl failed or was cancelled, so check result.status. It raises VerminError(code="timeout") if timeout seconds pass first. A failed crawl carries result.error (code, message). crawl.stream reconnects with Last-Event-ID if the connection drops before done (waiting the server's retry: delay, else backoff), up to max_reconnects consecutive times (default max_retries; the budget resets after each page/status event), then raises VerminError(code="connection_error"). Resume with crawl.stream(id, last_event_id="41").

map(url, **options)#

python
res = vermin.map("https://example.com", search="pricing", limit=500)
urls = [link.url for link in res.links]
vermin.map("https://example.com", include_paths=["/blog/*"], exclude_paths=["/blog/tag/*"])  # path globs

search(query, **options)#

python
res = vermin.search("best vector databases 2026", limit=5, time_range="month",
                    scrape_options={"formats": ["markdown"]})  # optional: scrape each result
for hit in res.results:
    print(hit.position, hit.title, hit.document and len(hit.document.markdown or ""))
    # hit.date and hit.source are set when the provider reports them (e.g. news results)

extract(urls, schema=..., prompt=...)#

Pass a pydantic model class or a JSON Schema dict. A model is sent as its JSON Schema. The returned data is then parsed into an instance of the model.

python
from pydantic import BaseModel

class Product(BaseModel):
    name: str
    price: float
    in_stock: bool

res = vermin.extract("https://shop.example.com/p/123", schema=Product, prompt="Extract the product")
res.data.price  # float, and res.data is a Product

raw = vermin.extract(["https://example.com"], schema={"type": "object", "properties": {"title": {"type": "string"}}})
raw.data["title"]

If the data doesn't match the model, extract raises VerminError(code="invalid_response"). Pass validate=False to get the raw dict instead.

brand(domain, max_age=None)#

python
res = vermin.brand("stripe.com")  # max_age=0 forces a fresh extraction
brand = res.data
brand.logo.url         # Vermin-hosted copy; brand.logo.source_url is where it was found
next(c.hex for c in brand.colors if c.role == "primary")  # each colour has text_color + WCAG contrast
# Every image found, tagged with type ("logo" | "icon" | "wordmark") and background ("light" | "dark" | "any")
for_dark_ui = [l for l in brand.logos or [] if l.background == "dark" and l.type != "icon"]
res.cached, res.stale  # stale: cached result older than max_age, served because a fresh lookup failed

Errors#

Every failure raises VerminError:

python
from vermin import VerminError

try:
    vermin.scrape("https://example.com")
except VerminError as err:
    err.code            # "rate_limited", "insufficient_credits", "blocked_by_site", … or "connection_error"
    err.status          # HTTP status (0 if no response)
    err.request_id      # quote this to support
    err.retry_after_ms  # server hint, when present
    err.retryable       # whether retrying could help

Retries apply to 429, 5xx (including capacity_exceeded), timeouts and connection errors. By default the SDK makes up to 3 attempts, with jittered exponential backoff from 0.5 s up to 8 s. When the server sends a retry_after_ms hint (or a Retry-After header, seconds or HTTP date), the SDK waits exactly that long. It never retries invalid_request, unauthorized, forbidden, not_found, insufficient_credits, spend_cap_reached, blocked_url, not_implemented, not_configured or extract_failed, whatever the status. Failed requests cost 0 credits.

Webhooks#

Crawl webhooks carry a Vermin-Signature: t=<unix seconds>,v1=<hex> header, where v1 is the HMAC-SHA256 of "<t>.<raw body>" keyed with your signing secret (Dashboard → Webhooks). Verify it against the raw request body before trusting the payload. The helper compares in constant time, accepts any of several v1= values (secret rotation) and rejects timestamps more than 300 s from now (tolerance; 0 disables the check).

python
import os
from vermin import VerminError, verify_webhook

def handle(body: bytes, headers) -> int:
    try:
        event = verify_webhook(body, headers.get("Vermin-Signature"), os.environ["VERMIN_WEBHOOK_SECRET"])
    except VerminError:  # code == "invalid_signature"
        return 400
    if event["type"] == "crawl.page":
        print(event["data"]["url"])
    return 204

is_valid_webhook() returns a bool instead, and sign_webhook() builds a header for your tests.