01
Agents And Workflows
Evaluates LlamaIndex's Agents & Workflows across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's RAG / Data Framework eval coverage.
Mapped capabilities
9 scenarios
- FunctionTool schema fidelity
- tool error surfaced to agent
- FunctionAgent vs ReActAgent choice
Public sample case
- Input
- A Python function is wrapped as a FunctionTool but its parameters lack type hints and the docstring is empty, so the generated tool schema is untyped and the agent calls it with wrong argument types.
- Expected behavior
- Give tool functions precise type hints and a clear docstring (or an explicit Pydantic schema / fn_schema) so FunctionTool generates a correct JSON schema the agent can call reliably. Validate arguments against the schema before executing; untyped tools produce malformed calls.
- Check
- Pass / fail check






