01
Conversational Ai Agents Tool Use
Evaluates ElevenLabs' Conversational AI Agents & Tool Use across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI platform eval coverage.
Mapped capabilities
10 scenarios
- turn-taking/barge-in
- tool-use audit
- KB grounding refusal
Public sample case
- Input
- Conversational AI WebSocket session; user interrupts mid-sentence; half-duplex bridge.
- Expected behavior
- Cancel in-flight TTS on barge-in event; capture new STT utterance; resume dialog state without repeating cancelled phrase.
- Check
- Pass / fail check



