Benchmarks
Measured, not claimed.
Every claim on this site maps to a metric, and every metric has a launch gate. The same harness runs on every change to the scraping code and publishes a nightly scorecard, generated straight from the results. Anyone can say they're the best. We'd rather be caught in the act.
Scorecard coming soon. The rats are still running the maze.
The baseline run is in progress. We'll publish the full scorecard, raw results and the public part of the corpus together, including the strata where we don't win yet.
Labelled corpus
1,000+ URLs, each with gold labels: main-content text, title, key links, tables and code blocks, plus logo, colour and name for brand URLs. Versioned, with a hidden holdout set so we can't overfit.
Stratified
Scored per stratum, not just overall, so a win on easy marketing pages can't hide a loss on bot-protected or JS-heavy ones.
Gated
Nothing ships unless every gate passes. In CI, regressions against the baseline block merges.
Strata
- static & marketing
- SPAs & JS-heavy
- docs
- e-commerce
- news & blogs
- forums
- bot-protected
- PDFs & files
- non-English & RTL
- huge pages
- slow origins
Contenders
Same corpus, same scoring, same machine budget.
- Vermin
- Firecrawl
- Context.dev
- Readability
- Trafilatura
- Crawl4AI
- Jina Reader
Metrics and launch gates
| Area | Metric | Gate | Result |
|---|---|---|---|
| Reliability | Success rate per stratum | ≥ best competitor; no stratum > 3 pts worse | pending |
| Content fidelity | Token-level F1 vs gold main text | ≥ best competitor + 2 pts | pending |
| Boilerplate | Nav/footer/cookie removal precision | Precision ≥ 0.95 | pending |
| Markdown structure | Per-element F1 (headings, lists, tables, code, links) | ≥ best competitor | pending |
| Metadata | Title/description/lang/canonical accuracy | ≥ 98% | pending |
| Map | URL recall vs sitemap ∪ link graph | Recall ≥ best competitor | pending |
| Crawl | Coverage, duplicate rate, pages/min | Dupes < 2% | pending |
| Search | nDCG@10 / answer-hit rate | ≥ Firecrawl search | pending |
| Extract | Field accuracy; schema-valid rate | Schema-valid 100% | pending |
| Brand | Logo, colour ΔE, name, fonts | Logo ≥ 97%, colour ≥ 95% | pending |
| Latency | p50/p95 per endpoint and tier | p50 ≤ best competitor | pending |
| Price | $ per 1K successful pages | Lower than both competitors at every tier | pending |