Endpoints
Extract
Typed JSON from up to 25 URLs, validated against your JSON Schema. Valid or free.
Extract scrapes every URL you give it, then has a model fill in one merged JSON value that is validated against your schema. Every successful response conforms: no "almost JSON", no missing required keys. Give a schema, a prompt, or both. 10 credits per URL scraped successfully (24 when the page needed the premium proxy). URLs that fail are free and listed in sources. If no valid JSON can be produced, the request fails with extract_failed at 0 credits.
Example#
curl -X POST "https://api.vermin.dev/v1/extract" \
-H "Authorization: Bearer $VERMIN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"urls": ["https://example.com/pricing"],
"schema": {
"type": "object",
"properties": {
"plans": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": { "type": "string" },
"monthlyUsd": { "type": "number" }
},
"required": ["name", "monthlyUsd"]
}
}
},
"required": ["plans"]
},
"prompt": "Every plan with its monthly price in USD"
}'Request#
urlsstring (uri)[]requiredmin 1 itemsmax 25 itemsschemaobjectJSON Schema for the outputpromptstringmax 4000 charsscrapeOptionsScrapeOptions
Response#
{
"success": true,
"data": {
"plans": [
{ "name": "Starter", "monthlyUsd": 19 },
{ "name": "Pro", "monthlyUsd": 89 }
]
},
"sources": [{ "url": "https://example.com/pricing", "ok": true }],
"credits_used": 10,
"request_id": "req_01J9V7D4NP"
}Writing good schemas#
- Require only what must exist. Every
requiredfield the page doesn't have is a reason to fail. Optional fields come back when present. - Describe fields.
descriptionon a property is passed to the model:"monthlyUsd": { "type": "number", "description": "Price per month, USD, before tax" }. - Keep it flat-ish. Deeply nested schemas work, but a flat list of objects is faster and more accurate.
- Use
promptfor judgement calls: "ignore enterprise plans", "prices in USD, convert if needed".