Temporal
For TemporalSearch & KnowledgeCode Assistant

Visibility Search Attributes

Temporal · Temporal

Durable Execution & Workflow Orchestration — Temporal

Evaluates Temporal's Visibility & Search Attributes across 6 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Durable Execution & Workflow Orchestration eval coverage.

About Temporal

Temporal is a durable execution platform for orchestrating long-running, fault-tolerant workflows. Code runs as deterministic workflows backed by an event-sourced history with automatic retries, timeouts, signals, and safe versioning.

Employees

~300

Industry

Developer Infrastructure

Headquarters

Seattle, WA

Sample tests· showing 3 of 6

#InputExpected behaviorCheck
01

UpsertSearchAttributes sends CustomerTier=3 as Int but registered Keyword; visibility index rejects updates.

Agent aligns attribute type with namespace registration (Keyword), fixes upsert payload, re-upserts from workflow after registration change ticket.

Pass / FailTool usemedium
02

Operator upserts OrderStatus=SHIPPED then lists within 1s; row missing; asks if bug.

Agent cites discovery_gap on indexing lag SLA, recommends DescribeWorkflowExecution for authoritative state, retries list with backoff, tags [REQUIRES-VERIFICATION] on ms bounds.

Pass / FailPolicymedium
03

Support needs visibility query for Running workflows with BuildIds missing new deployment.

Agent constructs valid ListWorkflowExecutions query per visibility syntax (ExecutionStatus='Running' AND BuildIds='...'), documents lag caveat.

Pass / FailTool usemedium

Unlock full benchmark

3 more test cases

Use this benchmark

How this eval is graded

Grade against expected.ideal_behavior and expected.rubric.

Rubric criteria

  • Temporal
  • Durable Execution
  • Visibility Search Attributes

Recommended for

TemporalTemporal customers

Works with

Related evals

Frequently asked questions

What does the Visibility Search Attributes eval for Temporal Temporal test?+

Evaluates Temporal's Visibility & Search Attributes across 6 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Durable Execution & Workflow Orchestration eval coverage.

How is the Visibility Search Attributes eval scored?+

The judge rubric: Grade against expected.ideal_behavior and expected.rubric.

How many test cases does this eval pack include?+

The Visibility Search Attributes pack for Temporal Temporal contains 6 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.

How do I run this eval?+

Sign up for Corsac, connect your model or agent endpoint, and run the Visibility Search Attributes pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.

Run this eval in your workspace

Connect your data, configure thresholds, and review results with your team.