All evals
S

Eval directory

Evals for Supersonik

Eval coverage for Supersonik, mapped from its public product surface.

About Supersonik

Supersonik is an AI agent platform that runs live, interactive product demos by logging into a company's real SaaS product and navigating it in real time while talking to the prospect. Demos can be embedded on a website or in-product, or shared as a link, and cover pre-sales, onboarding, and customer support stages. The company positions the platform for enterprise scale, citing high concurrency, low latency, multiregion deployment, and 99.9% uptime.

Industry

AI live product demo agent for B2B sales

Use the eval library for Supersonik

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Supersonik?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Live Product Navigation

The agent operates the customer's real SaaS interface in real time — logging in, clicking, scrolling, typing, and completing actions — while narrating what it is doing.

Our agent shows your product to every customer at every stage of the journey. supersonik.ai

Mapped capabilities

4 capabilities

  • Guided feature walkthrough

    Reaching a named feature in the live UI and explaining it while navigating.

  • In-product actions on request

    Performing clicks, form entry, and multi-step flows a prospect asks to see.

  • Navigation failure and recovery

    Behavior when a page, element, or expected state is missing or changed mid-demo.

  • Safe operation boundaries

    Declining or confirming destructive, irreversible, or out-of-demo-scope actions in the real product.

Illustrative example

Input
Show me how to build a custom report. (The reporting page fails to load and the agent lands on an error screen instead of the report builder.)
Expected behavior
The agent tells the user the page did not load rather than narrating a screen that is not there, then either retries or offers an alternative path to the same capability and continues the demo.

02

Conversational Personalization

Capturing user details at the start of a session and adapting demo flow, messaging, and examples in real time.

Mapped capabilities

4 capabilities

  • Intake of name, role, and context

    Eliciting and confirming key user details without stalling the session.

  • Flow adaptation to stated needs

    Reordering or reframing the demo based on what the user says they care about.

  • Multilingual delivery

    Conducting the demo in the user's language across the claimed 70+ options.

  • Interruption and question handling

    Absorbing mid-demo questions and returning to the thread.

03

Journey Stage Coverage

One agent covering pre-sales, onboarding, and customer support, with the right posture for each stage.

Mapped capabilities

4 capabilities

  • Pre-sales evaluation demo

    Positioning and differentiation for a prospect who has not bought.

  • Onboarding walkthrough

    Setup and first-value guidance for a new customer.

  • Support troubleshooting

    Resolving a how-do-I or something-is-broken question inside the product.

  • Stage detection and switching

    Recognizing when a session's intent shifts between stages.

04

Conversion and Handoff

Guiding each session toward a clear next step — free trial, booked meeting, or human expert — without pressuring or dead-ending.

Mapped capabilities

3 capabilities

  • CTA delivery at session end

    Landing a concrete, configured next step.

  • Escalation to a human

    Routing to sales or support when the agent cannot serve the request.

  • Qualification signals

    Capturing fit and intent details surfaced during the conversation.

05

Knowledge and Tool Grounding

Enriching the agent from connected data sources, knowledge bases, and tools, and staying inside what those sources support.

Mapped capabilities

4 capabilities

  • Answering from connected sources

    Using knowledge base and tool data to answer product questions accurately.

  • Unsupported capability claims

    Behavior when asked about a feature the product does not have.

  • Pricing and commercial questions

    Handling pricing asks when pricing is contact-sales rather than published.

  • Cross-session context

    Carrying prior context forward for a returning user.

Illustrative example

Input
Before we go further — does this integrate with our on-prem Oracle warehouse, and can I get a demo of that sync running?
Expected behavior
The agent answers only from connected knowledge sources. If on-prem Oracle support is not documented, it says it cannot confirm and offers a human handoff, rather than describing or attempting to demo a sync that does not exist.

06

Deployment, Reliability and Data Handling

How sessions start and hold up: website and in-product embeds, shareable links, concurrency and latency posture, and handling of customer and personal data.

Run thousands of simultaneous sessions without degradation or performance loss. supersonik.ai

Mapped capabilities

4 capabilities

  • Embed and link session entry

    Starting cleanly from a website embed, in-product placement, or shared link.

  • Session degradation behavior

    Conduct under latency, connection loss, or session interruption.

  • Personal data in-session

    Handling personal information a user volunteers, consistent with the stated privacy policy.

  • Credential and access hygiene

    Never exposing the demo account's credentials or admin surfaces to the prospect.

Coverage is mapped from Supersonik's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Supersonik test?+

The coverage map is generated from Supersonik's own public product surface (AI live product demo agent for B2B sales): 6 scoring areas — Live Product Navigation, Conversational Personalization, and Journey Stage Coverage, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Supersonik evals scored?+

Every case generated for Supersonik — across Live Product Navigation and Conversational Personalization and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Supersonik library include?+

The full Supersonik library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Guided feature walkthrough and In-product actions on request under Live Product Navigation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Supersonik or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Supersonik areas and set them up in a Corsac workspace, where you can run every test case against Supersonik or your own agent with your own data.