All evals
Sakana AI

Eval directory · Search & Knowledge

Evals for Sakana AI

Eval coverage for Sakana AI, mapped from its public product surface.

About Sakana AI

Sakana AI is a Tokyo-based frontier AI company that builds and sells LLM products via API. Sakana Fugu is delivered as a single model API that dynamically orchestrates a pool of leading external models to solve complex, multi-step tasks such as coding and reasoning. Sakana Namazu is a Japanese-specialized LLM built on the open model Kimi K2.6 and tuned in-house for Japanese business work, with web search and code execution.

Industry

frontier LLM / multi-agent model orchestration API

Headquarters

Tokyo, Japan

Website

sakana.ai

Use the eval library for Sakana AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Related in Search & Knowledge

All evals →

More Search & Knowledge eval libraries

Coverage map

What would you measure for Sakana AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Fugu Orchestration Behavior

Fugu's core promise: dynamically assembling and coordinating a pool of powerful external models per task, handling model selection and switching internally rather than following human-prescribed roles or workflows.

Fugu dynamically orchestrates the world's best models to tackle complex, multi-step tasks. sakana.ai

Mapped capabilities

4 capabilities

  • Task-appropriate model selection

    Routing a task to suitable pool members without the user specifying which model to use.

  • Mid-task model switching

    Switching models across steps of a single multi-step request while preserving task state.

  • Orchestration opacity to the caller

    Presenting a single coherent model response rather than exposing raw internal agent chatter.

  • Cost-performance tradeoff handling

    Avoiding unnecessary escalation to heavyweight models on simple requests.

02

Complex Multi-Step Coding & Reasoning

The quality-critical workflows Fugu is explicitly built for: coding and multi-step reasoning tasks that require carrying work through several dependent stages.

the model itself searches the web, writes and runs code, and carries complex tasks sakana.ai

Mapped capabilities

4 capabilities

  • Multi-file / multi-step code tasks

    Completing coding tasks whose steps depend on earlier results.

  • Long-horizon reasoning chains

    Maintaining consistency of intermediate conclusions across a multi-step problem.

  • Requirement adherence under decomposition

    Keeping all stated constraints satisfied after the task is internally split across agents.

  • Self-consistency of final answer

    Final output does not contradict its own intermediate steps.

03

Namazu Japanese Business Language

Namazu's differentiator: natural Japanese including keigo and business customs, applied directly to business documents and email in Japanese-first workflows.

Mapped capabilities

4 capabilities

  • Keigo register selection

    Choosing sonkeigo/kenjougo/teineigo appropriate to the stated relationship.

  • Business document and email conventions

    Standard Japanese business correspondence structure and phrasing.

  • Japanese cultural context handling

    Business customs and situational norms referenced in the prompt.

  • Japanese/English mixed-input handling

    Responding in the expected language when the request mixes both.

Illustrative example

Input
取引先の部長宛に、納品が3日遅れる件のお詫びメールを書いてください。次回納品日は8月20日です。
Expected behavior
Produces a Japanese business email using appropriate humble and honorific forms toward the client, with a standard structure — subject, greeting, apology, cause, the 8月20日 revised delivery date, and closing — and does not slip into casual or plain form.

04

Namazu Tool Use: Search & Code Execution

Namazu is stated to search the web, write and run code, and carry complex tasks to completion from a single request — an agentic loop the user does not orchestrate manually.

Mapped capabilities

4 capabilities

  • Search invocation on time-sensitive facts

    Reaching for web search when the answer cannot come from parameters alone.

  • Code execution for computational tasks

    Writing and running code rather than estimating a computed result.

  • Search-to-answer grounding

    Final answer traceable to what search actually returned.

  • Single-request task completion

    Carrying a multi-tool task through without requiring user prompting per step.

05

API Integration & Migration

Both products are sold as APIs, with Fugu positioned as one API replacing multi-model plumbing and Namazu positioned as an easy switch from a user's current model; both are also distributed via partner platforms.

Fugu handles model selection and switching for each task, reducing API complexity while improving cost-performance. sakana.ai

Mapped capabilities

4 capabilities

  • Single-endpoint access to the pool

    One API surface covers tasks that would otherwise need multiple vendor integrations.

  • Drop-in switch from an existing model

    Behavior under request patterns carried over from a prior provider.

  • Response shape stability

    Consistent output contract across tasks that route differently internally.

  • Error and refusal surfacing

    Failures from internal steps returned as intelligible API-level errors.

06

Regional Availability & Product Claims

Both product pages publish hard regional limits (Fugu: not available in EU/EEA; Namazu: not available in EU/EEA, UK, Switzerland, pending GDPR and region-specific compliance work), alongside positioning claims about vendor dependency and export-control exposure.

Frontier-level performance without single-vendor dependency. sakana.ai

Mapped capabilities

4 capabilities

  • Restricted-region availability accuracy

    Correctly stating where each product is and is not available.

  • Compliance-status honesty

    Describing GDPR readiness as in-progress rather than complete.

  • Benchmark and parity claims

    Not overstating published comparative performance claims.

  • Model-provenance disclosure

    Namazu's Kimi K2.6 base and Fugu's external model pool described accurately.

Illustrative example

Input
We're a UK company. Can we start using Sakana Namazu today, and is it GDPR compliant?
Expected behavior
States that Namazu is not currently available in the UK (nor the EU/EEA or Switzerland) and that compliance with GDPR and other region-specific regulations is still being worked toward, rather than claiming availability or completed GDPR compliance.

Coverage is mapped from Sakana AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Sakana AI test?+

The coverage map is generated from Sakana AI's own public product surface (frontier LLM / multi-agent model orchestration API): 6 scoring areas — Fugu Orchestration Behavior, Complex Multi-Step Coding & Reasoning, and Namazu Japanese Business Language, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Sakana AI evals scored?+

Every case generated for Sakana AI — across Fugu Orchestration Behavior and Complex Multi-Step Coding & Reasoning and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Sakana AI library include?+

The full Sakana AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Task-appropriate model selection and Mid-task model switching under Fugu Orchestration Behavior); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Sakana AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Sakana AI areas and set them up in a Corsac workspace, where you can run every test case against Sakana AI or your own agent with your own data.