All evals
Firecrawl

Eval directory · AI Platform

Evals for Firecrawl

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Firecrawl AI products.

About Firecrawl

Firecrawl is a web-data API for AI — it turns websites into clean, LLM-ready markdown or structured data via scrape, crawl, map, search, and LLM-powered extract endpoints, with JS rendering, browser actions, and proxies. Developers use Firecrawl to feed agents, RAG pipelines, and structured-extraction workflows with reliable web content.

Employees

~30

Industry

Web Data / Scraping

Headquarters

San Francisco, CA

Use the eval library for Firecrawl

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Firecrawl?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Actions And Dynamic Pages

Evaluates Firecrawl's Actions & Dynamic Pages across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Web Data for AI eval coverage.

Mapped capabilities

9 scenarios

  • click action before capture
  • scroll for infinite scroll
  • write/input into form fields

Public sample case

Input
Content sits behind a 'Load more' button; a plain scrape captures only the initial page state without the expanded content.
Expected behavior
Use an actions sequence with a click action targeting the button (plus a wait) so the dynamic content loads before capture. Verify the post-action markdown includes the expanded content rather than the initial state.
Check
Pass / fail check

02

Auth Rate Limits Credits Webhooks

Evaluates Firecrawl's Auth, Rate Limits, Credits & Webhooks across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Web Data for AI eval coverage.

Mapped capabilities

9 scenarios

  • Bearer API key auth
  • 429 + Retry-After backoff
  • concurrency cap discipline

Public sample case

Input
Agent puts the Firecrawl API key in a query string or a custom X-API-Key header and gets 401 Unauthorized.
Expected behavior
Send the key as Authorization: Bearer <FIRECRAWL_API_KEY>. Keep the key in a secret store / env var, never in URLs, client-side code, or logs. Rotate on suspected exposure.
Check
Pass / fail check

03

Crawl Whole Site

Evaluates Firecrawl's Crawl (whole site) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Web Data for AI eval coverage.

Mapped capabilities

9 scenarios

  • async job id + status polling
  • limit and maxDepth bounding
  • includePaths / excludePaths regex

Public sample case

Input
Agent calls POST /v1/crawl and immediately tries to read pages from the response body, but the response only contains a job id and a status URL.
Expected behavior
Treat crawl as asynchronous: capture the returned job id, then poll GET /v1/crawl/{id} until status=completed (or failed/cancelled) before consuming data[]. Use exponential backoff on polling rather than a tight loop.
Check
Pass / fail check

04

Extract Llm Structured

Evaluates Firecrawl's Extract (LLM structured) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Web Data for AI eval coverage.

Mapped capabilities

9 scenarios

  • JSON Schema definition
  • prompt + schema interplay
  • multi-URL extraction

05

Map Url Discovery

Evaluates Firecrawl's Map (URL discovery) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Web Data for AI eval coverage.

Mapped capabilities

9 scenarios

  • map vs crawl for enumeration
  • search filter on map
  • limit on map results

06

Safety Legality And Governance

Evaluates Firecrawl's Safety, Legality & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Web Data for AI eval coverage.

Mapped capabilities

10 scenarios

  • robots.txt respect
  • ToS / scraping legality
  • PII in scraped content

07

Scrape Single Url

Evaluates Firecrawl's Scrape (single URL) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Web Data for AI eval coverage.

Mapped capabilities

9 scenarios

  • formats array selection
  • onlyMainContent boilerplate stripping
  • includeTags / excludeTags scoping

Frequently asked questions

What do the Corsac evals for Firecrawl test?+

Each eval pack tests Firecrawl's public product surface — including Actions And Dynamic Pages, Auth Rate Limits Credits Webhooks, and Crawl Whole Site — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Firecrawl evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Firecrawl cases — from Safety Legality And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Firecrawl library.

How many test cases does the Firecrawl library include?+

The Firecrawl eval library includes 73 graded test cases across 8 eval packs, the largest being Safety Legality And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Firecrawl or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Firecrawl packs — Actions And Dynamic Pages and Auth Rate Limits Credits Webhooks and the rest — against Firecrawl or your own agent with your own data.