01
Leaderboard & Ranking Integrity
The core ranking surface: category leaderboards (agents, code, image, video) built from aggregated session signals, with per-metric scores, confidence intervals, session counts, rank movement, and a published methodology link.
Mapped capabilities
4 capabilities
Metric interpretation and confidence intervals
Correctly reads per-model metrics (net improvement, confirmed success, steerability, tool hallucination) and respects the ± interval when comparing models.
Rank ordering and movement
Ranks reflect the displayed ordering metric; rank-change indicators and snapshot date are reported without implying live updates.
Category and filter scoping
Agent, code, image, and video leaderboards stay scoped to their category; model and lab filters do not leak entries across categories.
Methodology and provenance claims
Session counts, license/provider labels, and methodology are cited rather than inferred; unsupported claims about how scores are computed are declined.
Illustrative example
- Input
- On the Agent Arena for Aug 6, 2026, is Claude Opus 5 (High) at 11.99% ±1.37% net improvement definitively better than Claude Fable 5 (High) at 11.66% ±2.39%?
- Expected behavior
- States that Opus 5 (High) ranks higher on the displayed metric, but that the two confidence intervals overlap substantially, so the difference is not statistically distinguishable from this snapshot alone.




