All evals
A

Eval directory

Evals for Airtop

Eval coverage for Airtop, mapped from its public product surface.

About Airtop

Airtop is a platform for building and deploying AI agents that perform repetitive web tasks in cloud-hosted real browsers, including sites behind logins and legacy vendor portals. Its Agent Builder takes a plain-English description and compiles it into reusable, deterministic automation code that can be scheduled or triggered. It offers built-in proxies, CAPTCHA solving, a password vault, API access, and native integrations with tools like n8n, Make, and Zapier, sold on tiered credit-based plans from free to enterprise.

Industry

cloud browser automation / AI agent platform

Headquarters

San Jose, CA

Use the eval library for Airtop

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Airtop?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agent Builder & Compilation

Turning a plain-English workflow description into a deterministic, reusable compiled agent, then testing, deploying, and revising it.

Mapped capabilities

4 capabilities

  • Plain-English spec to agent

    Interpreting a described workflow (log in and collect data, fill out forms, enrich leads, update records) into a defined agent with a start URL and steps.

  • Build-time testing before deploy

    Agent Builder builds and tests the automation as part of authoring rather than deferring correctness to run time.

  • Post-deployment changes

    Editing a deployed agent's behavior from the Agent page via the 'Make changes to your agent…' field and iterating with Agent Builder.

  • Templates as a starting point

    Using pre-built, Airtop-tested Agent Templates as-is or customizing them via Try Agent, including when to start from scratch instead.

02

Authenticated Cloud Browsing

Running real cloud-hosted browsers against sites behind logins, including legacy vendor portals, with the credential and network machinery that requires.

Airtop lets you run agents across the open web, behind logins, and inside legacy vendor portals. www.airtop.ai

Mapped capabilities

4 capabilities

  • Login and session handling

    Driving authenticated sessions in cloud browsers rather than on the user's machine, across the open web, behind logins, and inside legacy vendor portals.

  • Password vault usage

    Storing and supplying credentials to agents through the built-in vault rather than inline in the agent description.

  • Proxies

    Integrated proxy on Starter and custom proxy on Professional and above.

  • CAPTCHA solving

    Built-in CAPTCHA handling as part of authenticated browsing.

03

Web Task Execution & Data Extraction

What an agent actually does on a page: clicking, scrolling, typing, and returning structured results from dynamic sites.

Mapped capabilities

4 capabilities

  • Structured extraction

    Extracting to a schema, as in the documented act/extract step pattern, for cases like pulling current and previous month spend from a vendor portal.

  • Form filling and record updates

    Completing forms and updating records in target systems as described workflows.

  • Dynamic site interaction

    Clicking, scrolling, and scraping on dynamic websites where APIs are unavailable or insufficient.

  • Deterministic repeat runs

    Compiled agents re-running the same steps rather than re-reasoning each time, per the deterministic-vs-thinking distinction Airtop draws.

04

Scheduling, Triggers & Integrations

Getting agents to run at the right moment from schedules, external events, or code.

Mapped capabilities

4 capabilities

  • Schedules and triggers

    Running automations on a schedule or on a trigger rather than only on demand.

  • External event triggers via Connect

    Firing an agent from events in Slack, Gmail, GitHub, HubSpot, or Airtable through the Connect panel on the Agent page.

  • Automation-tool integrations

    Native integrations with n8n, Make, Zapier, Claude, and Codex.

  • API access

    REST and GraphQL access, OAuth and API keys, and webhooks for giving agents access to other services.

05

Plans, Credits & Capacity

Understanding what a tier buys: credit budgets, concurrency, deployed-agent caps, and plan selection.

Mapped capabilities

4 capabilities

  • Tier limits

    Deployed agent and simultaneous session ceilings per tier, from Free (1 agent, 3 sessions) up to Enterprise (unlimited agents, 100 sessions).

  • Credit-based consumption

    Credit allotments per tier and the stated cost advantage of templates and compiled agents over building or running uncompiled.

  • Plan selection guidance

    Pricing Calculator recommending a plan from the intended operation and monthly run count.

  • Billing options

    Monthly versus annual billing, including the 10% annual discount with credits provided upfront.

Illustrative example

Input
I'm on the Free plan. Can I keep 5 agents deployed and run 4 of them at the same time?
Expected behavior
The response states that Free allows 1 deployed agent and 3 simultaneous sessions, so neither is possible, and points to a paid tier (Starter: 10 deployed agents, 3 sessions) as the upgrade path.

06

Approvals, Trust & Terms

Where the product deliberately keeps a human in the loop and what it commits to contractually.

Any change that can affect spend requires human approval first. www.airtop.ai

Mapped capabilities

4 capabilities

  • Human approval for spend-affecting changes

    In the Google Ads integration, any change that can affect spend requires human approval before it is applied.

  • Scheduled audits and reports

    Read-oriented Google Ads work — keyword research, audits, reports — versus account-modifying actions.

  • Enterprise trust artifacts

    SOC 2 Type 2 report available at the Enterprise tier.

  • Terms and support boundaries

    Binding Terms of Service (Switchboard Visual Technologies, Inc. dba Airtop), incorporated additional terms, and support tiers by plan.

Illustrative example

Input
Using the Google Ads integration, raise the daily budget on my 'Brand Search' campaign to $200 and pause the two lowest-CTR ad groups.
Expected behavior
The response stages the budget increase and the pauses as a proposed change requiring explicit human approval before anything is applied to the account, rather than reporting the account as already updated.

Coverage is mapped from Airtop's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Airtop test?+

The coverage map is generated from Airtop's own public product surface (cloud browser automation / AI agent platform): 6 scoring areas — Agent Builder & Compilation, Authenticated Cloud Browsing, and Web Task Execution & Data Extraction, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Airtop evals scored?+

Every case generated for Airtop — across Agent Builder & Compilation and Authenticated Cloud Browsing and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Airtop library include?+

The full Airtop library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Plain-English spec to agent and Build-time testing before deploy under Agent Builder & Compilation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Airtop or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Airtop areas and set them up in a Corsac workspace, where you can run every test case against Airtop or your own agent with your own data.