Skip to content

newWeb data API for AI agents & RAG

Nothing on the web is safe.

Gets into every page. Leaves with the data.

Scrape, crawl, map, search, extract and brand any site from one edge-native API. Clean, chunked, LLM-ready markdown for your agents. Blocked? Then it's on us: failed requests cost 0 credits.

1,000 free credits every month, plus 2,000 on signup. No card. No traps. Or poke it in the playground first.

~/agent · vermin scrapereplay
$ vermin scrape https://shop.example/p/42 -f markdown,chunks▸ fetch     GET 200 · 38 KB · AMS · 41 ms✕ quality   js_required (empty <div id=root>)↑ escalate  render · headless chromium▸ render    200 · hydrated · 612 ms✓ extract   3,412 tokens · 8 chunks · −41% junk tier render  credits_used 1  latency 703 msgot in through the window. left with the data.
● fetch → render → premiumyou only pay for the page it brings back
  • static & marketing
  • SPAs & JS-heavy
  • docs
  • e-commerce
  • news & blogs
  • forums
  • bot-protected
  • PDFs & files
  • non-English & RTL
  • huge pages
  • slow origins

01First bite

One request. Clean meat.

POST a URL, get back markdown, chunks and metadata, plus a receipt: which tier it took to get in and exactly how many credits_used. No mystery multipliers, no “it depends”.

npm i @vermin/sdk
curl -X POST "https://api.vermin.dev/v1/scrape" \
  -H "Authorization: Bearer $VERMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://example.com/pricing", "formats": ["markdown", "chunks"] }'
200 OK184 mstier fetchcredits_used 1
{
  "success": true,
  "data": {
    "url": "https://example.com/pricing",
    "markdown": "# Pricing\n\nSimple plans for every team...",
    "chunks": [
      { "index": 0, "headingPath": ["Pricing"], "tokens": 412 }
    ],
    "metadata": { "title": "Pricing", "language": "en", "statusCode": 200 },
    "tier": "fetch",
    "cached": false,
    "quality": { "score": 0.97, "issues": [], "escalations": [] }
  },
  "credits_used": 1,
  "latency_ms": 184,
  "request_id": "req_01J9V3KX7Q"
}

02Auto-escalation

It finds a way in.

You don't pick a mode or guess which pages need a browser. Vermin tries the cheapest way in first and escalates only when the page fights back. Every response shows you which door it used, and why.

  1. 01 · The front door1 credit

    fetch

    A plain HTTP fetch from the edge. Fast, cheap and good enough for most of the web. Quality-checked before it counts.

  2. 02 · The window+0, included

    render

    Empty SPA shell? It climbs in with a headless browser, waits for hydration and tries again. Same price.

  3. 03 · The pipes15 credits, opt-in

    premium

    Bot wall in the way and you've allowed it? It goes through the plumbing: premium proxies, only when the others fail.

Transparent by default. The quality report lists every issue it found (js_required, blocked, thin_content) and every escalation it made. Pin a tier with tier, or cap it with maxTier.

response.data.quality
{
  "score": 0.96,
  "issues": ["js_required"],
  "escalations": [
    { "from": "fetch", "to": "render", "reason": "js_required" }
  ]
}

04Billing

Pay only for what it drags back.

Other APIs bill you for the attempt. We bill for the catch. If Vermin comes back empty-handed, the request is free. Always, on every plan, without a support ticket.

0 credits
for 4xx, 5xx, timeouts, blocks and empty shells it couldn't crack
JS included
rendering costs the same as a plain fetch
Rollover
unused credits carry into next month on every paid plan
Spend caps
hard caps and alerts on every plan, so the pack never eats your budget
403 blocked_by_siteYou pay: nothing
{
  "success": false,
  "error": {
    "code": "blocked_by_site",
    "message": "403 at every tier you allowed"
  },
  "credits_used": 0
}
200 OKYou pay: 1 credit
{
  "success": true,
  "data": { "markdown": "# The good stuff…", "tier": "fetch" },
  "credits_used": 1
}

05Edge-native

Lives in the pipes.

Vermin runs inside Cloudflare's network, in 300+ cities. Your request lands a few milliseconds from where it started, and the fetch leaves from wherever the origin is closest. No single datacenter to rate-limit, no region to pick, nowhere to keep it out of.

300+
cities
0
regions to configure
1
API, everywhere
Stylised network of Vermin edge locations connected by pipesSJCIADGRULHRAMSJNBBOMSINNRTSYD
10 of 300+ burrows shown● all connected

06LLM-ready output

Eats the junk. Keeps the meat.

Nav bars, cookie walls, footers and “you may also like” get chewed off before your model ever sees them. On a crawl, Vermin learns what repeats across the whole site and strips it everywhere, then hands you token-counted chunks with their heading path. Straight into your vector store.

what the page sent−41% tokens

Home · Products · Pricing · Blog · Log in · Sign up

🍪 We value your privacy. Accept all / Manage

# Rat-proof storage bins

Our bins are made from 1.2 mm steel with a gasket lid…

## Sizes

| Size | Litres | Price |

You may also like: Mouse trap deluxe, Cheese (bulk)

© 2025 Example Inc · Terms · Privacy · Sitemap

kept 522 tokens

chewed off 361 tokens

data.chunks
[
  {
    "index": 0,
    "headingPath": ["Rat-proof storage bins"],
    "text": "Our bins are made from 1.2 mm steel…",
    "tokens": 318
  },
  {
    "index": 1,
    "headingPath": ["Rat-proof storage bins", "Sizes"],
    "text": "| Size | Litres | Price |\n|---|---|---|…",
    "tokens": 204
  }
]
  • markdown, html, links, metadata
  • chunks with headingPath + tokens
  • crawl-wide boilerplate removal
  • tables, code and lists kept intact

07Brand intelligence

It steals the logo too.

Give it a domain, get the logo, icon, colour palette (each with a readable text colour), name, fonts and socials. Cached globally, so the second lookup is nearly free.

5

credits per lookup, 1 when cached.
Half what Context.dev charges.

GET /v1/brand?domain=vermin.devcredits_used 5

Vermin

fonts Bricolage Grotesque, Inter, Geist Mono

socials github · x · discord

  • Acid#c8f031
  • Ink#0a0a09
  • Bone#f3f1e8
  • Blood#ff6b57

08SDKs & MCP

Speaks your language.

Hand-written, idiomatic clients over one OpenAPI spec. Yes, including Rust, Clojure and Elixir, which nobody else bothers with, plus PHP. Typed errors, retries and streaming crawls built in.

  • TypeScript
  • Python
  • Rust
  • Clojure
  • Elixir
  • PHP
  • Hosted MCP
// npm i @vermin/sdk

import { Vermin } from "@vermin/sdk";

const vermin = new Vermin({ apiKey: process.env.VERMIN_API_KEY });

const { data, creditsUsed } = await vermin.scrape({ url: "https://example.com" });
console.log(data.tier, data.markdown, creditsUsed);
Install guides for all six →

Or let your agent sniff it out.

The MCP server gives Claude, Cursor or any MCP client all six endpoints as tools. Billed like the API.

# remote

https://mcp.vermin.dev/mcp

# local stdio

npx vermin-mcp

09Price

Cheaper to feed.

Up to 2.5× more credits per dollar than Firecrawl and 1.7× more than Context.dev, at list price, tier for tier. Before you count the failed requests we don't charge for.

Full pricing

Tier 1 · Starter $19/mo

$ per 1K credits

  • Vermin$1.58
  • Firecrawl$3.80
  • Context.dev$2.53

12,000 credits · 2.4× vs Firecrawl

Tier 2 · Pro $89/mo

$ per 1K credits

  • Vermin$0.59
  • Firecrawl$0.99
  • Context.dev$0.79

150,000 credits · 1.7× vs Firecrawl

Tier 3 · Growth $249/mo

$ per 1K credits

  • Vermin$0.42
  • Firecrawl$0.80
  • Context.dev$0.60

600,000 credits · 1.9× vs Firecrawl

Tier 4 · Scale $449/mo

$ per 1K credits

  • Vermin$0.30
  • Firecrawl$0.75
  • Context.dev$0.50

1,500,000 credits · 2.5× vs Firecrawl

List prices in USD, monthly billing. Shorter bar is better.

10Benchmarks

Measured, not claimed.

Everyone says they're the best. We built a labelled, stratified corpus to prove it, or embarrass us in public. Vermin, Firecrawl, Context.dev and the open-source extractors, same pages, same scoring, hidden holdout. Including where we lose.

See the methodology
Launch gates that must pass before release
AreaLaunch gateStatus
Reliability≥ best competitor; no stratum > 3 pts worsemeasuring
Content fidelity≥ best competitor + 2 ptsmeasuring
BoilerplatePrecision ≥ 0.95measuring
Markdown structure≥ best competitormeasuring
Metadata≥ 98%measuring
MapRecall ≥ best competitormeasuring
CrawlDupes < 2%measuring

11FAQ

Questions from the skirting board.

Anything else? The field guide has the rest.

What counts as a failed request?

Anything that doesn't bring back usable content: blocks, 4xx and 5xx responses, timeouts, and pages that stayed empty at every tier you allowed. They all return credits_used: 0.

How is this different from Firecrawl or Context.dev?

More credits per dollar at every tier, failed requests are free, JS rendering is included, escalation is transparent, chunks and crawl-wide boilerplate removal are built in, brand lookups cost half, and it runs at the edge. And we publish benchmarks against both.

Does it respect robots.txt?

Yes, by default. It also blocks private and internal addresses (SSRF protection) and identifies itself honestly. It gets in everywhere it's allowed to.

Do unused credits roll over?

On every paid plan, unused monthly credits carry into the next month. Add hard spend caps and alerts so overage never surprises you.

Can my agent use it directly?

Point any MCP client at https://mcp.vermin.dev/mcp or run npx vermin-mcp. Every endpoint shows up as a tool.

What do I get for free?

1,000 credits every month plus 2,000 on signup, with every endpoint. No card required.

Why on earth is it called Vermin?

Because it gets in everywhere, works all night, multiplies on demand and can't be kept out. Also, the domain was available.

Release the vermin.

First successful scrape in under two minutes. 1,000 free credits every month, plus 2,000 on signup. It'll be in the walls by lunch.