Skip to content

SDKs

Clojure SDK

The official Vermin client for Clojure. This page is the SDK's README, rendered.

dev.vermin/vermin

Official Clojure client for Vermin, the web-data API for AI: scrape, crawl, map, search, extract, brand.

Zero heavy dependencies: the JDK 11+ java.net.http client and org.clojure/data.json. No core.async.

clojure
;; deps.edn
{:deps {dev.vermin/vermin {:mvn/version "0.1.0"}}}

Quick start#

clojure
(require '[vermin.core :as vermin])

(def client (vermin/client {:api-key "vmn_live_..."}))   ; or (vermin/client) with VERMIN_API_KEY set

(def page (vermin/scrape client {:url "https://example.com"
                                 :formats [:markdown :metadata]
                                 :only-main-content true}))

(-> page :data :markdown)          ;=> "# Example Domain ..."
(-> page :data :metadata :title)   ;=> "Example Domain"
(:credits-used page)               ;=> 1

Requests are plain maps with kebab-case keys; they're sent as the API's camelCase (:only-main-content → onlyMainContent). Keyword values are converted the same way (:raw-html → "rawHtml"). Responses come back as kebab-case keyword maps (:credits-used, :final-url, :raw-html, …); enum values stay strings ("completed").

User-owned data is never case-converted: :schema, :headers and webhook :metadata on the way out, and extract's :data on the way back.

Configuration#

clojure
(vermin/client
  {:api-key     "vmn_live_..."            ; default $VERMIN_API_KEY
   :base-url    "https://api.vermin.dev"  ; default $VERMIN_API_URL, then https://api.vermin.dev
   :timeout-ms  60000                     ; per request (not event streams), default 90000
   :max-retries 3                         ; default 2
   :retry-base-ms 500 :retry-max-ms 8000
   :http-client my-java-net-http-client}) ; optional

A client is an immutable map; create one and share it across threads.

Endpoints#

clojure
;; Scrape — a bare URL string works too
(vermin/scrape client "https://example.com")
(vermin/scrape client {:url "https://example.com"
                       :formats [:markdown :chunks]
                       :tier :render
                       :actions [{:type :click :selector "#accept"} {:type :wait :ms 500}]})

;; Map, optionally filtered by path globs
(->> (vermin/map client {:url "https://example.com" :search "pricing"}) :links (map :url))
(vermin/map client {:url "https://example.com" :include-paths ["/blog/*"] :exclude-paths ["/blog/tag/*"]})

;; Search, optionally scraping every hit (news-style hits carry :date and :source)
(vermin/search client {:query "clojure web scraping" :limit 5 :scrape-options {:formats [:markdown]}})

;; Extract structured data (your schema's keys are preserved)
(-> (vermin/extract client {:urls ["https://example.com/pricing"]
                            :schema {:type "object"
                                     :properties {:planNames {:type "array" :items {:type "string"}}}}
                            :prompt "List plan names"})
    :data :planNames)

;; Brand — :logo/:icon plus every candidate in :logos (each with :type, :background,
;; :source-url); colours carry :text-color and its WCAG :contrast
(-> (vermin/brand client "stripe.com") :data :colors first :hex)
(->> (vermin/brand client "stripe.com") :data :logos (filter #(= "dark" (:background %))))
(vermin/brand client {:domain "stripe.com" :max-age 0})   ; 0 forces a fresh extraction
;; :cached / :stale on the response say whether it came from (possibly expired) cache

Crawling#

clojure
(def job (vermin/start-crawl client {:url "https://docs.example.com" :limit 200 :max-depth 3}))

;; Lazy seq of every document — pages through cursors and polls while running
(doseq [doc (vermin/crawl-seq client (:id job) {:poll-ms 2000})]
  (println (:url doc)))

;; Live server-sent events (seqable + closeable; closes itself after :done)
(with-open [events (vermin/crawl-events client (:id job))]
  (doseq [{:keys [event data]} events]
    (case event
      :page   (println "scraped" (:url data))
      :status (println (:completed data) "/" (:total data))
      :done   (println "finished, credits:" (:credits-used data))
      nil)))

;; Block until finished, collecting every page into :data. A failed crawl is
;; returned (not thrown) with :status "failed" and :error {:code :message}.
(vermin/wait-crawl client (:id job) {:poll-ms 2000 :timeout-ms 600000})

;; One page, cancel, or start+wait in one go
(vermin/get-crawl client (:id job) {:cursor nil :limit 100})
(vermin/cancel-crawl client (:id job))
(vermin/crawl client {:url "https://example.com" :limit 10})

If the connection drops before :done, crawl-events reconnects with Last-Event-ID (waiting the server's retry: delay, else backoff), up to :max-reconnects consecutive times (default :max-retries; the budget resets after each page/status event), then throws a :connection-error. Resume a stream yourself with (vermin/crawl-events client id {:last-event-id "41"}).

Errors#

Failures throw ex-info; ex-data is structured:

clojure
(try
  (vermin/scrape client "https://example.com")
  (catch clojure.lang.ExceptionInfo e
    (let [{:keys [code status request-id message retry-after-ms credits-used]} (ex-data e)]
      (case code
        :insufficient-credits (println "top up!")
        :rate-limited         (println "slow down, retry in" retry-after-ms "ms")
        (println code status request-id message)))))
  • :type is :vermin/api-error (use vermin/api-error?), :vermin/config-error or :vermin/timeout (from wait-crawl).
  • :code is the API code as a kebab keyword (:rate-limited, :not-found, :capacity-exceeded, :not-configured, :extract-failed, …; all in vermin/api-error-codes), or :connection-error / :timeout for transport failures.
  • (vermin/retryable? e) says whether retrying could help.
  • :credits-used is always 0 — failed requests are free.

Webhooks#

Crawl webhooks carry a Vermin-Signature: t=<unix seconds>,v1=<hex> header, where v1 is the HMAC-SHA256 of "<t>.<raw body>" keyed with your signing secret (Dashboard → Webhooks). Verify it against the raw request body before trusting the payload. The helper compares in constant time, accepts any of several v1= values (secret rotation) and rejects timestamps more than 300 s from now (tolerance; 0 disables the check).

clojure
(require '[vermin.webhook :as webhook])

(defn handler [{:keys [body headers]}]
  (let [raw (slurp body)
        secret (System/getenv "VERMIN_WEBHOOK_SECRET")]
    (try
      (let [event (webhook/verify raw (get headers "vermin-signature") secret)]
        (println (:type event))
        {:status 204})
      (catch clojure.lang.ExceptionInfo _ {:status 400}))))

verify throws ex-info with {:code :invalid-signature}; valid? returns a boolean and sign builds a header for your tests.

Retries and idempotency#

  • 408, 429, 5xx (including :capacity-exceeded) and connection errors are retried up to :max-retries times with jittered exponential backoff (base * 2^n, capped, jittered in [d/2, d]). Codes that would fail again are never retried, whatever the status: :invalid-request, :unauthorized, :forbidden, :not-found, :not-implemented, :insufficient-credits, :spend-cap-reached, :blocked-url, :not-configured, :extract-failed.
  • A server retry_after_ms (or Retry-After header, seconds or HTTP date) is honoured as a floor, plus up to 10% jitter.
  • Every POST gets an auto-generated UUID Idempotency-Key, reused across retries. Pass :idempotency-key "..." in the request map to set your own.

Specs#

vermin.specs defines clojure.spec models for every request and response (:vermin/scrape-request, :vermin/document, :vermin/crawl-status, :vermin/brand-response, :vermin/error-data, …):

clojure
(require '[clojure.spec.alpha :as s] 'vermin.specs)
(s/valid? :vermin/crawl-request {:url "https://example.com" :limit 50})