All evals
Sendoso

Eval directory · Customer Support

Evals for Sendoso

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Sendoso AI products.

About Sendoso

Sendoso is the leading sending platform for B2B go-to-market teams, enabling personalized direct mail, gifting, and physical experiences at scale. Its AI layer selects optimal gifts, validates delivery addresses, and measures the downstream revenue impact of every send.

Employees

~500

Industry

Sales Engagement & Gifting

Headquarters

Phoenix, AZ

Use the eval library for Sendoso

All 36 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Sendoso?

6 areas · 36 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Smart Delivery Address Ranking V1

Eval for Sendoso Smart Delivery ranking candidate addresses and refusing overconfident sends when office, home, and hybrid-work signals disagree.

Mapped capabilities

6 scenarios

  • Address Selection
  • False-Positive Prevention
  • Confidence-Based Holdouts

Example criterion: Smart Delivery should prefer the most likely real-world destination, not simply the cleanest address in a validation database.

02

Smart Delivery Failure Detection V1

Eval for Sendoso Smart Delivery detecting silent wrong-address outcomes and distinguishing data, process, and logistics failures after a gift is marked delivered.

Mapped capabilities

6 scenarios

  • Silent Failure Detection
  • Failure Attribution
  • Instrumentation Gaps

Example criterion: Smart Delivery should flag likely silent failures, assign the probable failure mode, and recommend the next detection or recovery action.

03

Smart Delivery Preflight Holds V1

Eval for Sendoso Smart Delivery deciding when to pause a shipment, request confirmation, or use a fallback path before a gift is sent to the wrong place.

Mapped capabilities

6 scenarios

  • Confirmation Triggers
  • Fallback Routing
  • Operational Escalations

Example criterion: Smart Delivery should stop bad sends before they happen and only escalate when the shipment risk is material or time-sensitive.

04

Smart Send Interest Match V1

Eval for Sendoso Smart Send choosing gifts from public interest signals without overfitting noisy social activity or violating policy.

Mapped capabilities

6 scenarios

  • Interest-To-Gift Match
  • Signal Weighting
  • Safe Fallback Behavior

Example criterion: Smart Send should choose gifts that are explainable, policy-safe, and grounded in strong recipient signals rather than shallow social noise.

05

Smart Send Outcome Learning V1

Eval for Sendoso Smart Send interpreting weak post-send feedback without over-claiming that a gift recommendation succeeded or failed.

Mapped capabilities

6 scenarios

  • Weak Positive Signals
  • Failure Attribution
  • Next-Best Instrumentation

Example criterion: Smart Send should classify outcome confidence honestly and recommend better instrumentation when the current feedback loop is too weak to trust.

06

Smart Send Signal Guardrails V1

Eval for Sendoso Smart Send deciding when to abstain, downgrade confidence, or route to review because public signals are stale, sparse, or unsafe.

Mapped capabilities

6 scenarios

  • Sparse Signal Handling
  • Sensitive Inference Guardrails
  • Operator Review Routing

Example criterion: Smart Send should know when not to trust the signal graph and when a neutral gift or manual review is safer than guessing.

Frequently asked questions

What do the Corsac evals for Sendoso test?+

Each eval pack tests Sendoso's public product surface — including Smart Delivery Address Ranking V1, Smart Delivery Failure Detection V1, Smart Delivery Preflight Holds V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Sendoso evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Sendoso library include?+

The Sendoso eval library includes 36 graded test cases across 6 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Sendoso or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Sendoso or your own agent with your own data.