SDKs
Clojure SDK
The official Vermin client for Clojure. This page is the SDK's README, rendered.
Official Clojure client for Vermin, the web-data API for AI: scrape, crawl, map, search, extract, brand.
Zero heavy dependencies: the JDK 11+ java.net.http client and org.clojure/data.json. No core.async.
;; deps.edn
{:deps {dev.vermin/vermin {:mvn/version "0.1.0"}}}Quick start#
(require '[vermin.core :as vermin])
(def client (vermin/client {:api-key "vmn_live_..."})) ; or (vermin/client) with VERMIN_API_KEY set
(def page (vermin/scrape client {:url "https://example.com"
:formats [:markdown :metadata]
:only-main-content true}))
(-> page :data :markdown) ;=> "# Example Domain ..."
(-> page :data :metadata :title) ;=> "Example Domain"
(:credits-used page) ;=> 1Requests are plain maps with kebab-case keys; they're sent as the API's camelCase (:only-main-content → onlyMainContent). Keyword values are converted the same way (:raw-html → "rawHtml"). Responses come back as kebab-case keyword maps (:credits-used, :final-url, :raw-html, …); enum values stay strings ("completed").
User-owned data is never case-converted: :schema, :headers and webhook :metadata on the way out, and extract's :data on the way back.
Configuration#
(vermin/client
{:api-key "vmn_live_..." ; default $VERMIN_API_KEY
:base-url "https://api.vermin.dev" ; default $VERMIN_API_URL, then https://api.vermin.dev
:timeout-ms 60000 ; per request (not event streams), default 90000
:max-retries 3 ; default 2
:retry-base-ms 500 :retry-max-ms 8000
:http-client my-java-net-http-client}) ; optionalA client is an immutable map; create one and share it across threads.
Endpoints#
;; Scrape — a bare URL string works too
(vermin/scrape client "https://example.com")
(vermin/scrape client {:url "https://example.com"
:formats [:markdown :chunks]
:tier :render
:actions [{:type :click :selector "#accept"} {:type :wait :ms 500}]})
;; Map, optionally filtered by path globs
(->> (vermin/map client {:url "https://example.com" :search "pricing"}) :links (map :url))
(vermin/map client {:url "https://example.com" :include-paths ["/blog/*"] :exclude-paths ["/blog/tag/*"]})
;; Search, optionally scraping every hit (news-style hits carry :date and :source)
(vermin/search client {:query "clojure web scraping" :limit 5 :scrape-options {:formats [:markdown]}})
;; Extract structured data (your schema's keys are preserved)
(-> (vermin/extract client {:urls ["https://example.com/pricing"]
:schema {:type "object"
:properties {:planNames {:type "array" :items {:type "string"}}}}
:prompt "List plan names"})
:data :planNames)
;; Brand — :logo/:icon plus every candidate in :logos (each with :type, :background,
;; :source-url); colours carry :text-color and its WCAG :contrast
(-> (vermin/brand client "stripe.com") :data :colors first :hex)
(->> (vermin/brand client "stripe.com") :data :logos (filter #(= "dark" (:background %))))
(vermin/brand client {:domain "stripe.com" :max-age 0}) ; 0 forces a fresh extraction
;; :cached / :stale on the response say whether it came from (possibly expired) cacheCrawling#
(def job (vermin/start-crawl client {:url "https://docs.example.com" :limit 200 :max-depth 3}))
;; Lazy seq of every document — pages through cursors and polls while running
(doseq [doc (vermin/crawl-seq client (:id job) {:poll-ms 2000})]
(println (:url doc)))
;; Live server-sent events (seqable + closeable; closes itself after :done)
(with-open [events (vermin/crawl-events client (:id job))]
(doseq [{:keys [event data]} events]
(case event
:page (println "scraped" (:url data))
:status (println (:completed data) "/" (:total data))
:done (println "finished, credits:" (:credits-used data))
nil)))
;; Block until finished, collecting every page into :data. A failed crawl is
;; returned (not thrown) with :status "failed" and :error {:code :message}.
(vermin/wait-crawl client (:id job) {:poll-ms 2000 :timeout-ms 600000})
;; One page, cancel, or start+wait in one go
(vermin/get-crawl client (:id job) {:cursor nil :limit 100})
(vermin/cancel-crawl client (:id job))
(vermin/crawl client {:url "https://example.com" :limit 10})If the connection drops before :done, crawl-events reconnects with Last-Event-ID (waiting the server's retry: delay, else backoff), up to :max-reconnects consecutive times (default :max-retries; the budget resets after each page/status event), then throws a :connection-error. Resume a stream yourself with (vermin/crawl-events client id {:last-event-id "41"}).
Errors#
Failures throw ex-info; ex-data is structured:
(try
(vermin/scrape client "https://example.com")
(catch clojure.lang.ExceptionInfo e
(let [{:keys [code status request-id message retry-after-ms credits-used]} (ex-data e)]
(case code
:insufficient-credits (println "top up!")
:rate-limited (println "slow down, retry in" retry-after-ms "ms")
(println code status request-id message))))):typeis:vermin/api-error(usevermin/api-error?),:vermin/config-erroror:vermin/timeout(fromwait-crawl).:codeis the API code as a kebab keyword (:rate-limited,:not-found,:capacity-exceeded,:not-configured,:extract-failed, …; all invermin/api-error-codes), or:connection-error/:timeoutfor transport failures.(vermin/retryable? e)says whether retrying could help.:credits-usedis always 0 — failed requests are free.
Webhooks#
Crawl webhooks carry a Vermin-Signature: t=<unix seconds>,v1=<hex> header, where v1 is the HMAC-SHA256 of "<t>.<raw body>" keyed with your signing secret (Dashboard → Webhooks). Verify it against the raw request body before trusting the payload. The helper compares in constant time, accepts any of several v1= values (secret rotation) and rejects timestamps more than 300 s from now (tolerance; 0 disables the check).
(require '[vermin.webhook :as webhook])
(defn handler [{:keys [body headers]}]
(let [raw (slurp body)
secret (System/getenv "VERMIN_WEBHOOK_SECRET")]
(try
(let [event (webhook/verify raw (get headers "vermin-signature") secret)]
(println (:type event))
{:status 204})
(catch clojure.lang.ExceptionInfo _ {:status 400}))))verify throws ex-info with {:code :invalid-signature}; valid? returns a boolean and sign builds a header for your tests.
Retries and idempotency#
- 408, 429, 5xx (including
:capacity-exceeded) and connection errors are retried up to:max-retriestimes with jittered exponential backoff (base * 2^n, capped, jittered in[d/2, d]). Codes that would fail again are never retried, whatever the status::invalid-request,:unauthorized,:forbidden,:not-found,:not-implemented,:insufficient-credits,:spend-cap-reached,:blocked-url,:not-configured,:extract-failed. - A server
retry_after_ms(orRetry-Afterheader, seconds or HTTP date) is honoured as a floor, plus up to 10% jitter. - Every POST gets an auto-generated UUID
Idempotency-Key, reused across retries. Pass:idempotency-key "..."in the request map to set your own.
Specs#
vermin.specs defines clojure.spec models for every request and response (:vermin/scrape-request, :vermin/document, :vermin/crawl-status, :vermin/brand-response, :vermin/error-data, …):
(require '[clojure.spec.alpha :as s] 'vermin.specs)
(s/valid? :vermin/crawl-request {:url "https://example.com" :limit 50})