SDKs
Rust SDK
The official Vermin client for Rust. This page is the SDK's README, rendered.
Official async Rust client for Vermin, the web-data API for AI: scrape, crawl, map, search, extract, brand.
[dependencies]
vermin = "0.1"
tokio = { version = "1", features = ["macros", "rt-multi-thread"] }Built on reqwest with rustls (no OpenSSL), serde and tokio.
Quick start#
use vermin::{Client, Format, ScrapeRequest};
#[tokio::main]
async fn main() -> vermin::Result<()> {
let client = Client::new("vmn_live_..."); // or Client::from_env()?
let res = client
.scrape(ScrapeRequest::new("https://example.com").formats([Format::Markdown, Format::Metadata]))
.await?;
println!("{}", res.data.markdown.unwrap_or_default());
println!("credits used: {}", res.credits_used);
Ok(())
}Configuration#
use std::time::Duration;
let client = vermin::Client::builder()
.api_key("vmn_live_...") // default: $VERMIN_API_KEY
.base_url("https://api.vermin.dev") // default: $VERMIN_API_URL, then https://api.vermin.dev
.timeout(Duration::from_secs(60)) // per request (not applied to event streams)
.max_retries(3) // default 2
.retry_backoff(Duration::from_millis(500), Duration::from_secs(8))
.build()?;Client is cheap to clone (Arc inside) — create one and share it.
Endpoints#
use serde_json::json;
use vermin::*;
// Scrape with options
let page = client.scrape(ScrapeRequest {
url: "https://example.com".into(),
options: ScrapeOptions {
formats: Some(vec![Format::Markdown, Format::Chunks]),
tier: Some(Tier::Render),
actions: Some(vec![Action::click("#accept"), Action::wait(500)]),
..Default::default()
},
}).await?;
// Map (optionally filtered by path globs)
let map = client.map("https://example.com").await?;
for link in &map.links { println!("{}", link.url); }
let blog = client.map(MapRequest {
include_paths: Some(vec!["/blog/*".into()]),
exclude_paths: Some(vec!["/blog/tag/*".into()]),
..MapRequest::new("https://example.com")
}).await?;
// Search (optionally scraping each hit; hits carry `date` / `source` when the provider reports them)
let hits = client.search(SearchRequest { limit: Some(5), ..SearchRequest::new("rust web scraping") }).await?;
// Extract structured data
#[derive(serde::Deserialize)]
struct Pricing { plans: Vec<String> }
let out = client.extract(
ExtractRequest::new(["https://example.com/pricing"])
.schema(json!({ "type": "object", "properties": { "plans": { "type": "array", "items": { "type": "string" } } } }))
.prompt("List plan names"),
).await?;
let pricing: Pricing = out.data_as()?;
// Brand: `logo`/`icon` plus every candidate in `logos`, each with `kind` (logo/icon/wordmark),
// `background` (light/dark/any) and `source_url`; colours carry `text_color` + WCAG `contrast`
let brand = client.brand("stripe.com").await?;
println!("{:?} {:?}", brand.data.name, brand.data.colors.first().map(|c| &c.hex));
let for_dark_ui: Vec<_> = brand.data.logos.iter()
.filter(|l| l.background == Some(ImageBackground::Dark) && l.kind != Some(ImageType::Icon))
.collect();
// `stale`: a cached result older than max_age, served because a fresh lookup failedCrawling#
use std::time::Duration;
use futures_util::StreamExt;
use vermin::*;
let job = client.crawls().start(CrawlRequest::new("https://docs.example.com").limit(200)).await?;
// Live server-sent events
let mut events = client.crawls().events(&job.id).await?;
while let Some(ev) = events.next().await {
match ev? {
CrawlEvent::Page(doc) => println!("scraped {}", doc.url),
CrawlEvent::Status(s) => println!("{:?}/{:?}", s.completed, s.total),
CrawlEvent::Done(s) => println!("finished: {} credits", s.credits_used.unwrap_or(0)),
_ => {}
}
}
// …or poll until finished and collect every page
let done = client.crawls().wait(&job.id, WaitOptions {
poll_interval: Duration::from_secs(2),
timeout: Some(Duration::from_secs(600)),
..Default::default()
}).await?;
println!("{} pages", done.data.len());
// One page of results / cancel
let page = client.crawls().get_page(&job.id, CrawlPageParams { cursor: None, limit: Some(100) }).await?;
client.crawls().cancel(&job.id).await?;
// Shortcut: start + wait
let all = client.crawl(CrawlRequest::new("https://example.com"), WaitOptions::default()).await?;If the connection drops before done, events reconnects with Last-Event-ID (waiting the server's retry: delay, else backoff), up to max_reconnects consecutive times (default max_retries; the budget resets after each page/status event), then yields Error::StreamClosed. Resume or tune it with crawls().events_with(&id, EventsOptions { last_event_id: Some("41".into()), ..Default::default() }). wait returns a failed crawl (with error: Some(CrawlError { code, message })) rather than erroring.
Errors#
Every fallible call returns vermin::Result<T>. API failures are Error::Api(ApiError) with code, message, status, request_id, retry_after and credits_used (always 0 on failure):
match client.scrape("https://example.com").await {
Ok(res) => println!("{} credits", res.credits_used),
Err(vermin::Error::Api(e)) if e.code == vermin::ErrorCode::InsufficientCredits => eprintln!("top up!"),
Err(e) => eprintln!("{e} (status={:?}, request_id={:?})", e.status(), e.request_id()),
}ErrorCode covers every API code (ErrorCode::ALL, e.g. CapacityExceeded, NotConfigured, ExtractFailed); err.is_retryable() says whether retrying could help.
Webhooks#
Crawl webhooks carry a Vermin-Signature: t=<unix seconds>,v1=<hex> header, where v1 is the HMAC-SHA256 of "<t>.<raw body>" keyed with your signing secret (Dashboard → Webhooks). Verify it against the raw request body before trusting the payload. The helper compares in constant time, accepts any of several v1= values (secret rotation) and rejects timestamps more than 300 s from now (tolerance; 0 disables the check).
// body: the raw request bytes; header: the Vermin-Signature value
match vermin::verify_webhook(&body, header, &secret) {
Ok(event) => println!("{}", event["type"]),
Err(vermin::Error::InvalidSignature(why)) => eprintln!("rejected: {why}"),
Err(e) => return Err(e.into()),
}vermin::webhook::{verify_at, is_valid_at, sign} take an explicit tolerance and clock (handy in tests).
Retries and idempotency#
- Retries 408, 429, 5xx (including
capacity_exceeded) and transport errors up tomax_retriestimes with jittered exponential backoff (base * 2^n, capped, uniformly jittered in[d/2, d]). Codes that would fail again are never retried, whatever the status:invalid_request,unauthorized,forbidden,not_found,not_implemented,insufficient_credits,spend_cap_reached,blocked_url,not_configured,extract_failed. - When the server sends
retry_after_ms(or aRetry-Afterheader, seconds or HTTP date) the SDK waits at least that long, plus up to 10% jitter. - Every POST carries an auto-generated UUID
Idempotency-Key, reused across retries of the same call. Override withRequestOptions::idempotency_key("…")via the*_withmethods (scrape_with,map_with,crawls().start_with, …).