All evals
Microsoft AutoGen

Eval directory · AI Platform

Evals for Microsoft AutoGen

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Microsoft AutoGen AI products.

About Microsoft AutoGen

Microsoft is a global technology company and a leading cloud and AI provider. Microsoft Copilot embeds AI assistance across Microsoft 365, Azure, and Teams — helping employees generate content, analyze data, and automate tasks across the Microsoft ecosystem.

Employees

~221,000

Industry

Enterprise Software & Cloud

Headquarters

Redmond, WA

Use the eval library for Microsoft AutoGen

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Microsoft AutoGen?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Autogen Agent Definitions

Evaluates Microsoft AutoGen's Agent Definitions (AssistantAgent / UserProxyAgent) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • system_message scope
  • model_client wiring
  • reflect_on_tool_use=True semantics

Public sample case

Input
Operator builds an AssistantAgent and passes the task instructions only via the first user message in team.run(task=...), leaving system_message at its default.
Expected behavior
Put persistent role/persona/policy instructions in AssistantAgent(system_message=...) so they appear as the system turn on every model_client call. Use the task argument only for the per-run user instruction. Without system_message, the agent loses role anchoring on multi-turn loops because the mod…
Check
Pass / fail check

02

Autogen Autogen Studio And Workbench

Evaluates Microsoft AutoGen's AutoGen Studio & Workbench across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • autogenstudio ui port binding
  • gallery component versioning
  • session trace persistence

Public sample case

Input
Operator runs `autogenstudio ui --host 0.0.0.0 --port 8080` on a shared dev VM with no firewall.
Expected behavior
Bind Studio to 127.0.0.1 by default in shared environments — the UI has no built-in auth gate and exposing it to a shared network gives any host on the network agent-construction + code-execution access. Front with an SSH tunnel or reverse-proxy with auth before binding to 0.0.0.0.
Check
Pass / fail check

03

Autogen Code Execution

Evaluates Microsoft AutoGen's Code Execution (Docker / Local) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • DockerCommandLineCodeExecutor lifecycle
  • LocalCommandLineCodeExecutor risk
  • work_dir bind-mount scope

Public sample case

Input
Operator constructs DockerCommandLineCodeExecutor inline in team.run() without calling executor.start() / stop().
Expected behavior
DockerCommandLineCodeExecutor requires an explicit start() before first use and stop() at teardown (or use it as an async context manager). Without start(), the first execute_code_blocks() call either lazily starts a container per call (slow + leaky) or errors. Wrap with 'async with' so the contain…
Check
Pass / fail check

04

Autogen Model Clients And Providers

Evaluates Microsoft AutoGen's Model Clients & Providers across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • OpenAIChatCompletionClient model_info
  • AzureOpenAIChatCompletionClient auth
  • streaming via run_stream + model client

05

Autogen Multi Agent Teams

Evaluates Microsoft AutoGen's Multi-agent Teams (RoundRobin / Selector / Swarm / MagenticOne) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • RoundRobinGroupChat speaker order
  • SelectorGroupChat selector_prompt grounding
  • allow_repeated_speaker semantics

06

Autogen Safety And Governance

Evaluates Microsoft AutoGen's Safety & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

10 scenarios

  • sandbox escape posture
  • secrets pulled from env
  • infinite-loop cost guardrails

07

Autogen Termination Conditions

Evaluates Microsoft AutoGen's Termination Conditions across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • TextMentionTermination exact match
  • MaxMessageTermination accounting
  • TokenUsageTermination thresholds

08

Autogen Tool Use And Function Calling

Evaluates Microsoft AutoGen's Tool Use & Function Calling across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • FunctionTool from typed callable
  • async tool dispatch
  • tool failure surfaces to model

Frequently asked questions

What do the Corsac evals for Microsoft AutoGen test?+

Each eval pack tests Microsoft AutoGen's public product surface — including Autogen Agent Definitions, Autogen Autogen Studio And Workbench, and Autogen Code Execution — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Microsoft AutoGen evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Microsoft AutoGen cases — from Autogen Safety And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Microsoft AutoGen library.

How many test cases does the Microsoft AutoGen library include?+

The Microsoft AutoGen eval library includes 73 graded test cases across 8 eval packs, the largest being Autogen Safety And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Microsoft AutoGen or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Microsoft AutoGen packs — Autogen Agent Definitions and Autogen Autogen Studio And Workbench and the rest — against Microsoft AutoGen or your own agent with your own data.