All evals
U

Eval directory

Evals for Ushur

Eval coverage for Ushur, mapped from its public product surface.

About Ushur

Ushur is an AI agent-powered platform for customer experience automation, aimed at healthcare, insurance, and financial services enterprises. It lets teams build, customize, launch, and analyze AI agents that run proactive inbound and outbound engagement across SMS/RCS, voice, email, and web, including its Invisible App secure-session experience. The platform markets a "trust-native architecture" with built-in governance, observability, and auditability, plus prebuilt solutions such as email triage, payments, member/policyholder service, and FNOL claims intake.

Industry

agentic AI customer experience automation platform

Website

ushur.ai

Use the eval library for Ushur

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Ushur?

6 scoring areas · 20 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Multi-Channel Engagement

Agent-driven inbound and outbound conversations across the channels Ushur documents, including continuity when an interaction moves from a messaging thread into a richer secure session.

Deliver proactive inbound and outbound engagement across every customer interaction — with compliance and auditability embedded by design. ushur.ai

Mapped capabilities

4 capabilities

  • SMS & RCS conversations

    Two-way messaging threads for proactive outbound and inbound inquiry handling.

  • Voice engagement

    Voice agent interactions and guided voice experiences, including reducing outbound call volume.

  • Email engagement

    Inbound email handling as a first-class conversational channel.

  • Web widget and Invisible App handoff

    On-site conversational widget plus escalation to a no-login secure session for forms, uploads, and payments.

02

Agent Lifecycle & Authoring

The build → customize → launch → analyze path teams follow to stand up an agent, tailor it to their process and brand, release it, and read its performance.

Maintain real-time visibility into agent behavior, reasoning, and performance. ushur.ai

Mapped capabilities

4 capabilities

  • Agent build and workflow authoring

    Low-code composition of agent journeys and steps within the Agentic Experience Framework.

  • Customization to enterprise process

    Tailoring agent behavior, content, and branching to a specific line of business.

  • Launch and rollout

    Moving a configured agent into live customer-facing operation across selected channels.

  • Analytics and performance reporting

    Reading resolution rates, engagement outcomes, and other agent performance signals.

03

Trust, Governance & Guardrails

The trust-native controls Ushur positions as core rather than layered: policy enforcement inside agent execution, human oversight, and the security posture regulated buyers require.

Built-in governance embeds guardrails, policy enforcement, and human oversight directly into AI execution. ushur.ai

Mapped capabilities

3 capabilities

  • Policy enforcement and guardrails

    Constraints applied during agent execution to keep behavior inside stated policy.

  • Human oversight and escalation

    Inserting human review or handoff into agent execution paths.

  • Security and compliance posture

    Encryption and industry-standard adherence for regulated healthcare, insurance, and finserv data.

04

Observability & Auditability

Runtime visibility into what an agent did and why, plus the ability to reconstruct an interaction after the fact for regulators or internal review.

Replay conversations end to end and generate regulator-ready records on demand. ushur.ai

Mapped capabilities

3 capabilities

  • Real-time agent behavior visibility

    Monitoring live agent behavior, reasoning, and performance.

  • End-to-end conversation replay

    Reconstructing a full interaction from first touch to resolution.

  • Regulator-ready record generation

    Producing audit records on demand from completed interactions.

Illustrative example

Input
Compliance reviewer requests the full record for a member service conversation that spanned an SMS thread and an escalated Invisible App session.
Expected behavior
The platform returns one end-to-end record covering both the messaging turns and the secure-session steps, in order, with timestamps and the agent actions taken, in a form the reviewer can export as an audit record.

05

Prebuilt Solution Workflows

The packaged journeys Ushur ships for common regulated-industry processes, each with its own intake, verification, and completion expectations.

delivering a robust FNOL automation solution at a 90% cost reduction. ushur.ai

Mapped capabilities

4 capabilities

  • Email triage

    Classifying and routing high-volume inbound enterprise email.

  • Payments

    Collecting a payment inside a secure session without requiring login.

  • Member and policyholder service

    Healthcare member/patient service and insurance policyholder service inquiries.

  • FNOL claims intake

    Digital first notice of loss capture through to a system-ready P&C claim.

Illustrative example

Input
Policyholder texts: "I backed into a pole in the parking garage this morning, need to file a claim." No policy number or photos provided.
Expected behavior
The agent opens an FNOL journey, requests the missing required items — policy identification, loss date and location, damage description, and photo upload via a secure session — and does not mark the claim system-ready until they are captured.

06

Integrations & Ecosystem

How agents exchange data with the systems of record and partner tooling that surround an enterprise CX deployment.

Mapped capabilities

2 capabilities

  • API-based integration

    Programmatic connections between agents and enterprise systems.

  • Partner and ecosystem connectors

    Integrations delivered through Ushur's partner ecosystem.

Coverage is mapped from Ushur's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Ushur test?+

The coverage map is generated from Ushur's own public product surface (agentic AI customer experience automation platform): 6 scoring areas — Multi-Channel Engagement, Agent Lifecycle & Authoring, and Trust, Governance & Guardrails, and more — spanning 20 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Ushur evals scored?+

Every case generated for Ushur — across Multi-Channel Engagement and Agent Lifecycle & Authoring and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Ushur library include?+

The full Ushur library is built on request. The coverage map spans 6 areas and 20 capabilities (for example, SMS & RCS conversations and Voice engagement under Multi-Channel Engagement); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Ushur or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Ushur areas and set them up in a Corsac workspace, where you can run every test case against Ushur or your own agent with your own data.