Skip to content

SDKs

PHP SDK

The official Vermin client for PHP. This page is the SDK's README, rendered.

vermin/vermin-phpcomposer require vermin/vermin-php guzzlehttp/guzzle

Official PHP SDK for Vermin, the web-data API for AI: scrape, crawl, map, search, extract and brand.

  • One Vermin client, built on PSR-18. It uses Guzzle if it's installed, otherwise any discovered client, or one you inject.
  • Typed, readonly response objects (with toArray() / JsonSerializable) and backed enums that follow the OpenAPI contract.
  • Automatic retries with jittered backoff on 429, 5xx and network errors. They honour retry_after_ms / Retry-After.
  • Auto-generated Idempotency-Key on scrape and crawl, reused across retries.
  • creditsUsed and requestId on every result, plus a typed exception hierarchy.
  • Crawl polling, cursor pagination as generators, and server-sent-event streaming with automatic reconnects.
  • Webhook signature verification.

Requires PHP 8.1+.

bash
composer require vermin/vermin-php guzzlehttp/guzzle

Guzzle is optional. Any PSR-18 client with PSR-17 factories works (for example symfony/http-client + nyholm/psr7).

Quick start#

php
use Vermin\Vermin;
use Vermin\Enum\Format;

$vermin = new Vermin(); // reads VERMIN_API_KEY (and optionally VERMIN_API_URL)

$page = $vermin->scrape('https://example.com', ['formats' => [Format::Markdown, Format::Metadata]]);
echo $page->data->markdown, $page->data->metadata?->title, "\n";
echo "cost: {$page->creditsUsed} credits (request {$page->requestId})\n";

Options#

php
$vermin = new Vermin(
    apiKey: 'vmn_live_…',                // default: VERMIN_API_KEY
    baseUrl: 'https://api.vermin.dev',   // default: VERMIN_API_URL, then https://api.vermin.dev
    timeout: 60.0,                       // per-attempt HTTP timeout, seconds
    maxRetries: 2,                       // retries on 429 / 5xx / network errors
    headers: ['X-Team' => 'growth'],     // extra headers on every request
    httpClient: $psr18Client,            // bring your own PSR-18 client
    requestFactory: $psr17Factory,       // optional PSR-17 factories (discovered by default)
    streamFactory: $psr17Factory,
);

Request parameters use the API's camelCase names ('onlyMainContent' => true) and accept enums or their string values (Format::Markdown or 'markdown'). null values are left out.

Every method also takes a RequestOptions for a single call:

php
use Vermin\RequestOptions;

$vermin->scrape('https://example.com', options: new RequestOptions(
    timeout: 120.0,
    maxRetries: 0,
    idempotencyKey: 'order-42-scrape',
    headers: ['X-Trace' => 'abc'],
));

The timeout applies when the client is Guzzle (the default). With other PSR-18 clients, set timeouts on the client you inject.

API#

scrape($url, $params)#

php
use Vermin\Enum\{ActionType, Format, QualityIssue, Tier};

$res = $vermin->scrape('https://news.ycombinator.com', [
    'formats' => [Format::Markdown, Format::Links, Format::Chunks],
    'onlyMainContent' => true,
    'tier' => Tier::Auto,             // fetch → render → premium (if allowPremium)
    'chunkSize' => 1200,
    'actions' => [
        ['type' => ActionType::Click, 'selector' => '#more'],
        ['type' => ActionType::Wait, 'ms' => 1000],
    ],
]);
foreach ($res->data->chunks ?? [] as $chunk) {
    echo implode(' › ', $chunk->headingPath ?? []), " ({$chunk->tokens} tokens)\n";
}
$res->data->quality?->hasIssue(QualityIssue::JsRequired);

Crawl#

php
use Vermin\Enum\CrawlEventType;

$job = $vermin->crawl('https://docs.example.com', [
    'limit' => 200,
    'maxDepth' => 3,
    'scrapeOptions' => ['formats' => ['markdown']],
    'webhook' => ['url' => 'https://you.dev/hooks/vermin', 'events' => ['page', 'completed']],
]);

// Poll until finished and collect every page (follows `next` cursors):
$result = $vermin->waitForCrawl($job->id, pollInterval: 2.0, timeout: 600,
    onStatus: fn ($s) => print("{$s->completed} / {$s->total}\n"));
echo $result->status->value, ' ', count($result->data), ' ', $result->creditsUsed, "\n";

// …or stream pages as they arrive (server-sent events):
foreach ($vermin->streamCrawl($job->id) as $event) {
    match ($event->type) {
        CrawlEventType::Page => print($event->document()->url . "\n"),
        CrawlEventType::Done => print('finished: ' . $event->status()->status->value . "\n"),
        default => null,
    };
}

// Lower level:
$status = $vermin->getCrawl($job->id, limit: 50);
$page2 = $status->next ? $vermin->getCrawl($job->id, cursor: $status->next) : null;
foreach ($vermin->crawlDocuments($job->id) as $doc) {  // every document, across pages
    echo $doc->url, "\n";
}
foreach ($vermin->crawlPages($job->id) as $page) {     // one CrawlStatus per page
    echo count($page->data), "\n";
}
$vermin->cancelCrawl($job->id);

waitForCrawl returns the final status even if the crawl failed or was cancelled, so check $result->status (or $result->isDone()). It throws TimeoutException if timeout seconds pass first. A failed crawl carries $result->error (['code' => …, 'message' => …]). streamCrawl reconnects with Last-Event-ID if the connection drops before done (waiting the server's retry: delay, else backoff), up to maxReconnects consecutive times (default maxRetries; the budget resets after each event), then throws NetworkException. Resume with streamCrawl($id, lastEventId: '41'). Breaking out of the foreach closes the connection.

map($url, $params)#

php
$res = $vermin->map('https://example.com', ['search' => 'pricing', 'limit' => 500]);
$urls = $res->urls(); // or array_map(fn ($l) => $l->url, $res->links)

// Only some paths (globs)
$vermin->map('https://example.com', ['includePaths' => ['/blog/*'], 'excludePaths' => ['/blog/tag/*']]);

search($query, $params)#

php
use Vermin\Enum\TimeRange;

$res = $vermin->search('best vector databases 2026', [
    'limit' => 5,
    'timeRange' => TimeRange::Month,
    'scrapeOptions' => ['formats' => ['markdown']], // optional: scrape each result
]);
foreach ($res->results as $hit) {
    echo $hit->position, ' ', $hit->title, ' ', strlen($hit->document?->markdown ?? ''), "\n";
    // $hit->date and $hit->source are set when the provider reports them (e.g. news results)
}

extract($urls, $schema, $prompt)#

Pass a JSON Schema as an array (or any JsonSerializable). The schema is sent as-is, and data comes back as decoded JSON (arrays).

php
$res = $vermin->extract('https://shop.example.com/p/123', schema: [
    'type' => 'object',
    'properties' => [
        'name' => ['type' => 'string'],
        'price' => ['type' => 'number'],
        'in_stock' => ['type' => 'boolean'],
    ],
    'required' => ['name', 'price'],
], prompt: 'Extract the product');

$res->data['price'];
foreach ($res->sources ?? [] as $source) {
    echo $source->url, "\n";
}

brand($domain, $maxAge)#

php
use Vermin\Enum\{ColorRole, ImageBackground, ImageType};

$res = $vermin->brand('stripe.com');
$brand = $res->data;
$brand->logo?->url;          // Vermin-hosted copy; ->sourceUrl is where it was found
$brand->color(ColorRole::Primary)?->hex;
$brand->color(ColorRole::Primary)?->contrast; // WCAG ratio of ->textColor on it

// Every image found, tagged with its type and the background it suits
$forDarkUi = array_filter($brand->logos ?? [], fn ($l) => $l->background === ImageBackground::Dark && $l->type !== ImageType::Icon);

$res->cached; // served from cache
$res->stale;  // cached result older than maxAge, served because a fresh lookup failed

$vermin->brand('stripe.com', maxAge: 0); // force a fresh extraction

Every response object has toArray() (the API's wire format) and is JsonSerializable. Enum fields hold a backed enum, or the raw string if the API returns a value this SDK version doesn't know about yet.

Errors#

Every failure throws a Vermin\Exception\VerminException subclass:

php
use Vermin\Exception\{InsufficientCreditsException, RateLimitException, VerminException};

try {
    $vermin->scrape('https://example.com');
} catch (RateLimitException $e) {
    sleep((int) ceil($e->retryAfter() ?? 1));
} catch (InsufficientCreditsException $e) {
    // top up at https://vermin.dev/dashboard/billing
} catch (VerminException $e) {
    $e->errorCode;     // ErrorCode::BlockedBySite, … (or the raw string for new codes)
    $e->status;        // HTTP status (0 if no response)
    $e->requestId;     // quote this to support
    $e->retryAfterMs;  // server hint, when present
    $e->creditsUsed;   // credits charged, when the API reports it
    $e->isRetryable(); // whether retrying could help
}
ExceptionError codes
ValidationExceptioninvalid_request (400)
AuthenticationExceptionunauthorized (401)
PermissionDeniedExceptionforbidden (403)
NotFoundExceptionnot_found (404)
InsufficientCreditsExceptioninsufficient_credits, spend_cap_reached (402)
RateLimitExceptionrate_limited (429), capacity_exceeded (503)
FetchExceptionfetch_failed, blocked_by_site, blocked_url, unsupported_content
TimeoutExceptiontimeout (API or client-side)
NotImplementedExceptionnot_implemented, not_configured (501)
ExtractionExceptionextract_failed (422: pages fetched, extraction failed)
ServerExceptioninternal and other 5xx
NetworkExceptionconnection_error (no response)
InvalidResponseExceptioninvalid_response (the body didn't match the contract)
WebhookSignatureExceptioninvalid_signature
MissingApiKeyExceptionmissing_api_key (thrown by the constructor when no key is configured)

Retries apply to 429, 5xx, timeouts and connection errors. By default the SDK makes up to 3 attempts, with jittered exponential backoff from 0.5 s up to 8 s. When the server sends a retry_after_ms hint or a Retry-After header, the SDK waits exactly that long. It never retries invalid_request, unauthorized, forbidden, not_found, insufficient_credits, spend_cap_reached, blocked_url, not_implemented, not_configured or extract_failed. Failed requests cost 0 credits.

Webhooks#

Crawl webhooks are signed with HMAC-SHA256 using your endpoint's secret. The Vermin-Signature header looks like t=1767225600,v1=5257a869…, where v1 is the hex HMAC of "{t}.{raw body}". Verify it against the raw request body before trusting the payload:

php
use Vermin\Webhook;
use Vermin\Exception\WebhookSignatureException;

$payload = file_get_contents('php://input');
$header = $_SERVER['HTTP_VERMIN_SIGNATURE'] ?? '';

try {
    $event = Webhook::verify($payload, $header, getenv('VERMIN_WEBHOOK_SECRET'), tolerance: 300);
} catch (WebhookSignatureException $e) {
    http_response_code(400);
    exit;
}

// $event is the decoded JSON payload (array<string, mixed>)

verify compares signatures in constant time, accepts any of several v1= values (for secret rotation), and rejects timestamps more than tolerance seconds from now (0 disables that check). Webhook::isValid() returns a bool instead, and Webhook::sign() builds a header for your tests.

Configuration#

SettingConstructor argumentEnvironment
API keyapiKeyVERMIN_API_KEY
Base URLbaseUrlVERMIN_API_URL (default https://api.vermin.dev)
Timeouttimeout (seconds, default 60)—
RetriesmaxRetries (default 2)—

Requests send Authorization: Bearer <key> and User-Agent: vermin-php/<version> PHP/<major.minor>.