All evals
T

Eval directory

Evals for TinyFish

Eval coverage for TinyFish, mapped from its public product surface.

About TinyFish

TinyFish is an API platform that gives AI agents access to the live web through four products: Search (structured web search), Fetch (URL to clean content), Agent (multi-step web automation), and Browser (cloud browser sessions). It is built on Mako, a web-native model positioned for completing real workflows on live websites, including dynamic and authenticated flows. Plans range from free/pay-as-you-go credits to an Enterprise tier with custom credits, SLA, and on-premise options.

Industry

AI web agent and web data infrastructure APIs

Use the eval library for TinyFish

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for TinyFish?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Search API

Fast, structured web search that returns fresh, browser-rendered results from the live web rather than a cache, aimed at agents tracking fast-moving sources.

Search the web, fetch clean content from any URL, run a real browser, and deploy agents. www.tinyfish.ai

Mapped capabilities

4 capabilities

  • Freshness and non-cached results

    Explaining that results come from live, browser-rendered search and are never served from a cache.

  • Structured JSON result shape

    Describing the returned fields (title, url, snippet) and how an agent consumes them.

  • Dynamic-source coverage

    When Search is the right tool for pages and sources that change too fast for traditional search.

  • Rate limits and free access

    Search costs 0 credits per request on every plan, under per-plan request-per-minute limits.

02

Fetch API

Turns any URL into clean, agent-ready page context, loading pages in a real browser before extraction.

Search and Fetch APIs are free. Agent and Browser APIs consume credits. www.tinyfish.ai

Mapped capabilities

4 capabilities

  • Real-browser rendering

    JavaScript-rendered content and dynamic layouts come through without the caller managing browser infrastructure.

  • Chrome and boilerplate removal

    Navigation, cookie banners, ads, and repeated footers fall away while text, headings, lists, and structure remain.

  • Output format selection

    Returning markdown, HTML, or JSON depending on what the downstream agent needs.

  • Search-to-Fetch handoff

    Moving from discovered URLs to clean destination-page evidence as a single grounded path.

Illustrative example

Input
Can Fetch pull the enterprise rate limit out of a docs page that renders its table in JavaScript, and what formats can it hand back?
Expected behavior
Confirms Fetch loads the page in a real browser so JavaScript-rendered content is captured, strips navigation and boilerplate, and returns clean markdown, HTML, or JSON. Notes that multi-step or logged-in navigation belongs to Web Agent, not Fetch.

03

Web Agent

Goal-driven multi-step automation on live websites, for work that cannot be solved by reading a single page.

Mapped capabilities

4 capabilities

  • Goal-to-completion execution

    Accepting a stated goal and navigating, filling, and submitting without a hand-written script.

  • Authenticated and session-based flows

    Working inside dashboards, accounts, and portals that require a logged-in session.

  • Adapting to changing paths

    Handling workflows where the path is dynamic or unknown in advance.

  • Structured result extraction

    Returning results the calling agent can consume, with browser work kept outside the agent's context.

04

Browser API

Cloud browser sessions with the anti-bot, stealth, and proxy infrastructure included rather than self-managed.

Mapped capabilities

4 capabilities

  • Session lifecycle and duration

    Sessions billed per 4 minutes against a 60-minute cap.

  • Concurrency limits

    How many browser sessions or agent runs can be in flight on a given plan.

  • Stealth, anti-bot, and proxy

    Capabilities included on every plan rather than priced separately.

  • Vault and Profile

    Included credential and profile handling for repeat access to the same sites.

05

Mako Model Positioning

The web-native model underneath the APIs, positioned for reliable completion of real workflows on live sites at a fraction of frontier-model cost.

Reliable at scale, at a fraction of frontier-model cost. www.tinyfish.ai

Mapped capabilities

4 capabilities

  • Web-native vs general-purpose framing

    Why a model built to operate the live web is presented as different from a general frontier model.

  • Portal-first workflows

    Targeting sites that expose no API, where the portal itself is the product.

  • Reliability at scale

    Positioning around repeated, high-volume runs rather than one-off demos.

  • Cost posture

    Claimed economics relative to frontier-model alternatives.

06

Plans, Credits, and Enterprise

How usage converts to cost across pay-as-you-go, Starter, Pro, and Enterprise, plus the terms that matter to a buyer.

Both are included on every plan and use no credits. www.tinyfish.ai

Mapped capabilities

4 capabilities

  • Credit metering by API

    Search and Fetch at 0 credits; Agent per step; Browser per 4-minute block.

  • Plan selection and overage

    Included monthly credits and per-credit overage across PAYG, Starter, and Pro.

  • Failed-run billing

    Failed runs cost $0 on every plan.

  • Enterprise tier terms

    Custom credits, SLA, on-premise, compliance, research mode, and priority support.

Illustrative example

Input
I'm on the Starter plan. Which of your four APIs consume credits, how is Browser metered, and what am I charged when an Agent run fails?
Expected behavior
States that Search and Fetch consume no credits on any plan, Agent bills 1 credit per step, and Browser bills 1 credit per 4 minutes up to a 60-minute cap. States that failed runs cost $0.

Coverage is mapped from TinyFish's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for TinyFish test?+

The coverage map is generated from TinyFish's own public product surface (AI web agent and web data infrastructure APIs): 6 scoring areas — Search API, Fetch API, and Web Agent, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the TinyFish evals scored?+

Every case generated for TinyFish — across Search API and Fetch API and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the TinyFish library include?+

The full TinyFish library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Freshness and non-cached results and Structured JSON result shape under Search API); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against TinyFish or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped TinyFish areas and set them up in a Corsac workspace, where you can run every test case against TinyFish or your own agent with your own data.