01
Evaluation Runs Eval Correctness
Mapped capabilities
30 scenarios
- dataset_empty
Public sample case
- Input
- CI pipeline agent is asked to run the nightly regression eval for the "checkout-flow" prompt before promoting a new prompt version. The pipeline config still references dataset_id ds_8a2f00112233checkout_v3, but a recent dataset …
- Expected behavior
- Agent inspects the run summary, sees total rows scored = 0, and reports explicitly that no rows were evaluated (e.g. 'The run scored 0 of 0 rows — dataset_id ds_8a2f... resolved to an empty set, likely stale/renamed. I cannot determine pass/fail and recommend NOT promoting v12 until this is fixed.'…
- Check
- Pass / fail check






