Skip to content

SDKs

Rust SDK

The official Vermin client for Rust. This page is the SDK's README, rendered.

vermin

Official async Rust client for Vermin, the web-data API for AI: scrape, crawl, map, search, extract, brand.

toml
[dependencies]
vermin = "0.1"
tokio = { version = "1", features = ["macros", "rt-multi-thread"] }

Built on reqwest with rustls (no OpenSSL), serde and tokio.

Quick start#

rust
use vermin::{Client, Format, ScrapeRequest};

#[tokio::main]
async fn main() -> vermin::Result<()> {
    let client = Client::new("vmn_live_...");            // or Client::from_env()?

    let res = client
        .scrape(ScrapeRequest::new("https://example.com").formats([Format::Markdown, Format::Metadata]))
        .await?;
    println!("{}", res.data.markdown.unwrap_or_default());
    println!("credits used: {}", res.credits_used);
    Ok(())
}

Configuration#

rust
use std::time::Duration;

let client = vermin::Client::builder()
    .api_key("vmn_live_...")                  // default: $VERMIN_API_KEY
    .base_url("https://api.vermin.dev")       // default: $VERMIN_API_URL, then https://api.vermin.dev
    .timeout(Duration::from_secs(60))         // per request (not applied to event streams)
    .max_retries(3)                           // default 2
    .retry_backoff(Duration::from_millis(500), Duration::from_secs(8))
    .build()?;

Client is cheap to clone (Arc inside) — create one and share it.

Endpoints#

rust
use serde_json::json;
use vermin::*;
// Scrape with options
let page = client.scrape(ScrapeRequest {
    url: "https://example.com".into(),
    options: ScrapeOptions {
        formats: Some(vec![Format::Markdown, Format::Chunks]),
        tier: Some(Tier::Render),
        actions: Some(vec![Action::click("#accept"), Action::wait(500)]),
        ..Default::default()
    },
}).await?;

// Map (optionally filtered by path globs)
let map = client.map("https://example.com").await?;
for link in &map.links { println!("{}", link.url); }
let blog = client.map(MapRequest {
    include_paths: Some(vec!["/blog/*".into()]),
    exclude_paths: Some(vec!["/blog/tag/*".into()]),
    ..MapRequest::new("https://example.com")
}).await?;

// Search (optionally scraping each hit; hits carry `date` / `source` when the provider reports them)
let hits = client.search(SearchRequest { limit: Some(5), ..SearchRequest::new("rust web scraping") }).await?;

// Extract structured data
#[derive(serde::Deserialize)]
struct Pricing { plans: Vec<String> }
let out = client.extract(
    ExtractRequest::new(["https://example.com/pricing"])
        .schema(json!({ "type": "object", "properties": { "plans": { "type": "array", "items": { "type": "string" } } } }))
        .prompt("List plan names"),
).await?;
let pricing: Pricing = out.data_as()?;

// Brand: `logo`/`icon` plus every candidate in `logos`, each with `kind` (logo/icon/wordmark),
// `background` (light/dark/any) and `source_url`; colours carry `text_color` + WCAG `contrast`
let brand = client.brand("stripe.com").await?;
println!("{:?} {:?}", brand.data.name, brand.data.colors.first().map(|c| &c.hex));
let for_dark_ui: Vec<_> = brand.data.logos.iter()
    .filter(|l| l.background == Some(ImageBackground::Dark) && l.kind != Some(ImageType::Icon))
    .collect();
// `stale`: a cached result older than max_age, served because a fresh lookup failed

Crawling#

rust
use std::time::Duration;
use futures_util::StreamExt;
use vermin::*;
let job = client.crawls().start(CrawlRequest::new("https://docs.example.com").limit(200)).await?;

// Live server-sent events
let mut events = client.crawls().events(&job.id).await?;
while let Some(ev) = events.next().await {
    match ev? {
        CrawlEvent::Page(doc) => println!("scraped {}", doc.url),
        CrawlEvent::Status(s) => println!("{:?}/{:?}", s.completed, s.total),
        CrawlEvent::Done(s) => println!("finished: {} credits", s.credits_used.unwrap_or(0)),
        _ => {}
    }
}

// …or poll until finished and collect every page
let done = client.crawls().wait(&job.id, WaitOptions {
    poll_interval: Duration::from_secs(2),
    timeout: Some(Duration::from_secs(600)),
    ..Default::default()
}).await?;
println!("{} pages", done.data.len());

// One page of results / cancel
let page = client.crawls().get_page(&job.id, CrawlPageParams { cursor: None, limit: Some(100) }).await?;
client.crawls().cancel(&job.id).await?;

// Shortcut: start + wait
let all = client.crawl(CrawlRequest::new("https://example.com"), WaitOptions::default()).await?;

If the connection drops before done, events reconnects with Last-Event-ID (waiting the server's retry: delay, else backoff), up to max_reconnects consecutive times (default max_retries; the budget resets after each page/status event), then yields Error::StreamClosed. Resume or tune it with crawls().events_with(&id, EventsOptions { last_event_id: Some("41".into()), ..Default::default() }). wait returns a failed crawl (with error: Some(CrawlError { code, message })) rather than erroring.

Errors#

Every fallible call returns vermin::Result<T>. API failures are Error::Api(ApiError) with code, message, status, request_id, retry_after and credits_used (always 0 on failure):

rust
match client.scrape("https://example.com").await {
    Ok(res) => println!("{} credits", res.credits_used),
    Err(vermin::Error::Api(e)) if e.code == vermin::ErrorCode::InsufficientCredits => eprintln!("top up!"),
    Err(e) => eprintln!("{e} (status={:?}, request_id={:?})", e.status(), e.request_id()),
}

ErrorCode covers every API code (ErrorCode::ALL, e.g. CapacityExceeded, NotConfigured, ExtractFailed); err.is_retryable() says whether retrying could help.

Webhooks#

Crawl webhooks carry a Vermin-Signature: t=<unix seconds>,v1=<hex> header, where v1 is the HMAC-SHA256 of "<t>.<raw body>" keyed with your signing secret (Dashboard → Webhooks). Verify it against the raw request body before trusting the payload. The helper compares in constant time, accepts any of several v1= values (secret rotation) and rejects timestamps more than 300 s from now (tolerance; 0 disables the check).

rust
// body: the raw request bytes; header: the Vermin-Signature value
match vermin::verify_webhook(&body, header, &secret) {
    Ok(event) => println!("{}", event["type"]),
    Err(vermin::Error::InvalidSignature(why)) => eprintln!("rejected: {why}"),
    Err(e) => return Err(e.into()),
}

vermin::webhook::{verify_at, is_valid_at, sign} take an explicit tolerance and clock (handy in tests).

Retries and idempotency#

  • Retries 408, 429, 5xx (including capacity_exceeded) and transport errors up to max_retries times with jittered exponential backoff (base * 2^n, capped, uniformly jittered in [d/2, d]). Codes that would fail again are never retried, whatever the status: invalid_request, unauthorized, forbidden, not_found, not_implemented, insufficient_credits, spend_cap_reached, blocked_url, not_configured, extract_failed.
  • When the server sends retry_after_ms (or a Retry-After header, seconds or HTTP date) the SDK waits at least that long, plus up to 10% jitter.
  • Every POST carries an auto-generated UUID Idempotency-Key, reused across retries of the same call. Override with RequestOptions::idempotency_key("…") via the *_with methods (scrape_with, map_with, crawls().start_with, …).