All evals
Black Forest Labs

Eval directory

Evals for Black Forest Labs

Eval coverage for Black Forest Labs, mapped from its public product surface.

About Black Forest Labs

Black Forest Labs is a frontier AI lab that builds the FLUX family of generative visual models, with FLUX 3 as a single multimodal foundation model spanning image, video, audio, and action-prediction for robotics. The models are offered as a hosted playground and production API, and as open weights that customers can license, fine-tune, and deploy on their own infrastructure. Pricing is usage-based per generation, with tiered open-weights licenses and custom enterprise agreements.

Employees

90-person

Industry

multimodal generative AI models (image, video, audio)

Headquarters

Freiburg, Germany and San Francisco

Website

bfl.ai

Use the eval library for Black Forest Labs

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Black Forest Labs?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Image Generation

Still-image generation grounded in the real world: stylistic range, prompt adherence for complex multi-part instructions, and the accurate in-image text rendering the site calls out as a headline capability.

Mapped capabilities

4 capabilities

  • Text rendering accuracy in images

    Requested strings appear in the generated image spelled correctly and legibly, including longer phrases and mixed casing.

  • Complex prompt adherence

    Multi-clause prompts specifying subject, count, spatial relation, and setting are all honored in a single generation.

  • Style breadth and control

    Named or described visual styles are followed without collapsing to a single default cinematic look.

  • Real-world grounding

    Objects, materials, and physical relationships in the image remain plausible rather than internally contradictory.

Illustrative example

Input
Generate a photorealistic image of a corner bakery at dusk. The awning sign reads exactly: FLOUR & SALT — EST. 1994. No other visible text anywhere in the scene.
Expected behavior
The rendered sign shows the requested string character-for-character, including the ampersand, em dash, and year, with correct spelling and casing. No additional invented signage or garbled lettering appears elsewhere in the frame.

02

Video Generation

Text-, image-, and keyframe-initiated video, including video-to-video, with the duration, resolution tier, and multi-shot behavior advertised for FLUX 3.

One multimodal model for Image Video Audio Action-Prediction bfl.ai

Mapped capabilities

4 capabilities

  • Text / image / keyframe conditioning

    Each documented entry point produces video that respects the supplied starting condition.

  • Duration and resolution modes

    Draft, HD, and FHD selections and clip lengths up to the stated 20-second ceiling behave as specified.

  • Multi-shot single take

    A prompt describing several shots yields distinct shots within one generation rather than one continuous take.

  • Video-to-video transformation

    Source motion and structure are preserved while the requested transformation is applied.

03

Native Audio

Audio generated jointly with frames rather than dubbed afterward: multilingual speech, sound effects, and ambience that line up with on-screen events.

native audio and up to 20 second clips in a single generation bfl.ai

Mapped capabilities

4 capabilities

  • Audio-visual synchronization

    Sound events correspond in time to the visual events that would cause them.

  • Multilingual speech

    Requested spoken language is produced as specified alongside the generated frames.

  • Effects and ambience

    Environmental and incidental audio matches the depicted setting.

  • Optional audio toggle

    Audio is present or absent according to the request, without degrading the video track.

04

Action-Prediction for Robotics

The FLUX-mimic surface described in the blog: visual observations plus text instructions in, predicted physical outcome and robot control actions out, unifying perception, simulation, and execution.

Run FLUX models on your own infrastructure. Full control over deployment, fine-tuning, and customization. bfl.ai

Mapped capabilities

4 capabilities

  • Observation + instruction intake

    Visual input paired with a natural-language instruction is accepted as a single conditioned request.

  • Visual outcome prediction

    The predicted end state reflects the instructed manipulation.

  • Control action output

    Predictions are emitted as robot control actions, not only as rendered frames.

  • Physical plausibility

    Predicted motion respects contact, weight, and cause-and-effect rather than producing impossible transitions.

05

API & Playground Integration

The hosted delivery path — playground for experimentation, documented API and dashboard for production — described as built to handle production workloads at any scale.

Mapped capabilities

4 capabilities

  • API request/response contract

    Documented parameters are accepted and responses conform to the documented shape.

  • Playground-to-API parity

    Settings exercised in the playground map to equivalent API parameters.

  • Usage-based billing transparency

    Reported generation cost matches the published per-unit rates for the selected mode and duration.

  • Production-scale behavior

    Concurrent and sustained request patterns are handled per the stated scale positioning.

Illustrative example

Input
Request a text-to-video generation at FHD, 5 seconds, and read back the quoted price from the pricing calculator and the API usage response for the same configuration.
Expected behavior
Both surfaces quote the same total, computed as the published per-second rate multiplied by the requested duration — $0.17 × 5 = $0.85 — with no hidden per-request or seat fee added, consistent with the stated pay-as-you-go model.

06

Open Weights & Licensing

Self-hosted delivery: downloadable models with fine-tuning and LoRA rights, tiered by model access, monthly image volume, domains, and licensed users, up through enterprise and synthetic-data terms.

Simple-to-integrate API to access the latest and most powerful FLUX models. bfl.ai

Mapped capabilities

4 capabilities

  • Tier entitlement boundaries

    Model access, monthly image volume, domain count, and seat count match the purchased tier.

  • Fine-tuning and LoRA rights

    Customization permitted by the tier is exercisable and correctly documented.

  • Self-hosted deployment

    Published weights run on customer infrastructure, including the on-device and Klein variants named in the material.

  • Enterprise commitments

    Stated compliance and support claims (SOC 2, ISO 27001, SLA, custom volume) are represented accurately in contract-facing surfaces.

Coverage is mapped from Black Forest Labs's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Black Forest Labs test?+

The coverage map is generated from Black Forest Labs's own public product surface (multimodal generative AI models (image, video, audio)): 6 scoring areas — Image Generation, Video Generation, and Native Audio, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Black Forest Labs evals scored?+

Every case generated for Black Forest Labs — across Image Generation and Video Generation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Black Forest Labs library include?+

The full Black Forest Labs library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Text rendering accuracy in images and Complex prompt adherence under Image Generation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Black Forest Labs or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Black Forest Labs areas and set them up in a Corsac workspace, where you can run every test case against Black Forest Labs or your own agent with your own data.