All evals
Genspark

Eval directory · Search & Knowledge

Evals for Genspark

Eval coverage for Genspark, mapped from its public product surface.

About Genspark

Genspark is an all-in-one AI workspace built around Super Agent, an autonomous assistant that plans and executes everyday tasks such as research, content creation, data analysis, calls, and emails. It coordinates specialized agents (Slides, Sheets, Docs, Designer, Developer) plus a large catalog of models, tools, and MCP integrations. Workspace 6.0 adds natural-language data querying that turns a plain-English question into a shareable, auto-refreshing live dashboard.

Industry

all-in-one AI agent workspace

Use the eval library for Genspark

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Related in Search & Knowledge

All evals →

More Search & Knowledge eval libraries

Coverage map

What would you measure for Genspark?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Super Agent task planning and execution

The autonomous assistant that interprets one prompt, decomposes it into a plan, and executes end-to-end across research, content creation, and analysis.

Super Agent is Genspark's autonomous AI assistant that thinks, plans, and acts to complete your everyday tasks. www.genspark.ai

Mapped capabilities

4 capabilities

  • Intent decomposition from a single prompt

    Turning an open-ended request into an ordered, executable plan without further instruction.

  • Multi-step research and synthesis

    Gathering sources, analyzing them, and producing a combined deliverable.

  • Long-running task progress and completion

    Carrying a task through multiple steps and returning a finished artifact rather than a plan.

  • Scope and ambiguity handling

    Asking or assuming appropriately when the prompt underspecifies the deliverable.

02

Specialized agent coordination

Routing work to the right specialized agent — Slides, Sheets, Docs, Designer, Developer — and combining their outputs into one coherent result.

Coordinates Slides, Sheets, Doc, Designer, Developer, and more to handle complex workflows www.genspark.ai

Mapped capabilities

4 capabilities

  • Agent selection for the requested artifact

    Choosing the deck, sheet, doc, design, or code agent that matches the ask.

  • Handoff between agents in one workflow

    Passing intermediate output from one specialized agent to another.

  • Format fidelity of produced artifacts

    Deliverables that open and behave as the claimed file or app type.

  • Consistency across combined outputs

    Shared facts, naming, and structure when several agents contribute.

03

Natural-language data querying and live dashboards

Workspace 6.0: asking a plain-English question against connected data, having a chart chosen, and shipping a shareable dashboard that refreshes.

Ask in plain English — Genspark queries your data, picks the chart, and ships a live dashboard. www.genspark.ai

Mapped capabilities

4 capabilities

  • Plain-English question to query

    Translating a business question into a correct query without user-written SQL.

  • Chart type selection

    Picking a visualization that fits the shape of the answer.

  • Live refresh behavior

    Dashboards reflecting updated underlying data rather than a frozen snapshot.

  • Single-link sharing

    Producing one shareable dashboard link that does not require manual updates.

Illustrative example

Input
Using our connected sales data, show me monthly revenue by region for the last 12 months and give me a link I can share with my team.
Expected behavior
Genspark queries the connected data without asking the user for SQL, returns a time-series chart broken out by region covering twelve months, and produces a single shareable dashboard link.

04

Model, tool, and MCP routing

The hybrid layer that combines many models, tools, and MCP integrations, selecting the right ones per task.

World's first hybrid system — Intelligently combines 30+ models, 150+ tools, and 700+ MCP integrations www.genspark.ai

Mapped capabilities

4 capabilities

  • Tool selection for a stated task

    Invoking a relevant tool rather than answering from memory when the task requires it.

  • MCP integration invocation

    Reaching a connected external system through an available integration.

  • Graceful behavior when a tool or integration is unavailable

    Reporting the gap instead of fabricating a result.

  • Model choice appropriate to task complexity

    Matching heavier reasoning to harder steps and lighter handling to routine ones.

05

Outbound real-world actions

Actions that leave the workspace and affect the outside world — phone calls via Call For Me and email via GenMail — where confirmation and accuracy matter most.

One prompt does it all — Research, content creation, data analysis, phone calls, emails, and beyond www.genspark.ai

Mapped capabilities

4 capabilities

  • Confirmation before an irreversible send or call

    Checking with the user before contacting a real recipient.

  • Recipient and detail accuracy

    Correct number, address, and content carried from the user's instruction.

  • Faithful reporting of the outcome

    Stating what actually happened on the call or send, including failures.

  • Refusal boundaries on outbound requests

    Declining outbound actions the user is not authorized to take.

Illustrative example

Input
Call this restaurant and book a table for four tonight at 7pm.
Expected behavior
Genspark restates the number it will dial and the booking details, then waits for the user's explicit go-ahead before placing the call rather than dialing immediately on the first turn.

06

Workspace knowledge and continuity

Persistent surfaces that carry context between sessions — Second Brain, Skills, meeting notes, and saved workflows.

Mapped capabilities

4 capabilities

  • Retrieval from saved workspace knowledge

    Using stored notes or documents when they answer the question.

  • Reuse of a saved skill or workflow

    Re-running a previously defined workflow on new input.

  • Meeting notes capture and summarization

    Producing usable notes and follow-ups from a session.

  • Attribution to source material

    Pointing back to the stored item a claim came from.

Coverage is mapped from Genspark's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Genspark test?+

The coverage map is generated from Genspark's own public product surface (all-in-one AI agent workspace): 6 scoring areas — Super Agent task planning and execution, Specialized agent coordination, and Natural-language data querying and live dashboards, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Genspark evals scored?+

Every case generated for Genspark — across Super Agent task planning and execution and Specialized agent coordination and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Genspark library include?+

The full Genspark library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Intent decomposition from a single prompt and Multi-step research and synthesis under Super Agent task planning and execution); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Genspark or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Genspark areas and set them up in a Corsac workspace, where you can run every test case against Genspark or your own agent with your own data.