All evals
StaffScheduleCare

Eval directory · Document Agents

Evals for StaffScheduleCare

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for StaffScheduleCare AI products.

About StaffScheduleCare

StaffScheduleCare is a workforce-management platform for scheduling, payroll, compliance, and staffing operations in long-term care and hospitals.

Industry

Healthcare Workforce-Management Software

Headquarters

Newmarket, Ontario

Use the eval library for StaffScheduleCare

All 11 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Document Agents

All evals →

More Document Agents eval libraries

Coverage map

What would you measure for StaffScheduleCare?

1 area · 11 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ssc Parser Regression V1

Regression eval library for SSC's deterministic XLSX intake-form parser. Each fixture is a small synthetic workbook designed to lock in one specific parser code path tied to a known/resolved bug — when the parser regresses on that path, the fixture name tells you which fix broke. This dataset contains the 11 synthetic regression fixtures only. SSC's local CI runs an extended dataset that also includes their 5 real client fixtures (kept in the SSC repo for data-isolation reasons). EXECUTION MODEL — EXTERNAL ONLY (until Phase 2): This spec is designed to run via the SSC SDK in the SSC repo's CI, which invokes the TypeScript parser locally and reports per-example results to Corsac. The dataset shape ({input: {fixture_path}, expected: {data, errors, warnings}}) reflects that workflow — it does NOT match the doc_agent runner's input/output contract (input.document/input.instruction → {result: ...}). Triggering "Run" from the Corsac UI will currently fail with structural mismatches; that's expected. Phase 2 work fixes this two ways: (a) adds a 'parser_regression' category whose runner skips server-side execution, and (b) ships the external-run endpoint that accepts pushed results from SSC's CI. Tracked in issues/phase-2-ssc-pilot-readiness.md (2-09, 2-17) and issues/phase-1.5b-autogen-hardening.md.

Mapped capabilities

11 scenarios

    Frequently asked questions

    What do the Corsac evals for StaffScheduleCare test?+

    Each eval pack tests StaffScheduleCare's public product surface — including Ssc Parser Regression V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

    How are the StaffScheduleCare evals scored?+

    Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

    How many test cases does the StaffScheduleCare library include?+

    The StaffScheduleCare eval library includes 11 graded test cases across 1 eval pack. Each case defines an input, expected behavior, and pass/fail criteria.

    How do I run these evals against StaffScheduleCare or my own agent?+

    Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against StaffScheduleCare or your own agent with your own data.