All evals
P

Eval directory

Evals for Probabl

Eval coverage for Probabl, mapped from its public product surface.

About Probabl

Skore is an AI "data scientist" agent from probabl, the company behind scikit-learn, that builds and improves machine learning models while enforcing scientific rigor, explainability, and reproducibility. It positions itself as a "rigor layer" over agent-generated ML pipelines, working across AI providers, ML frameworks, and compute environments. It is offered as open-source software (`pip install skore`) alongside commercial SaaS, private cloud, and on-prem deployments.

Industry

AI data science agent / ML methodology platform

Headquarters

Paris, France

Use the eval library for Probabl

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Probabl?

6 scoring areas · 22 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agentic modeling workflow

The core promise: an agent teammate that drafts, runs, and improves a machine learning pipeline on a real problem, from a loaded dataset through an improved model.

From the scikit-learn company, a new teammate that ships machine learning models with the scientific rigor www.probabl.ai

Mapped capabilities

4 capabilities

  • Problem loading and task framing

    Selecting a problem such as customer churn and establishing what is being predicted before modeling starts.

  • Pipeline drafting

    Producing a runnable baseline pipeline for the loaded problem.

  • Iterative model improvement

    Proposing and applying changes that improve a model, with the change attributed to a result.

  • Handling unavailable problems

    Behavior for surfaces marked coming soon, such as credit default risk and employee salaries.

02

Scientific rigor and methodology

The differentiator claimed by the scikit-learn maintainers: statistical methodology baked into agent-generated work rather than bolted on after a result exists.

skore-agent - methodology for agents, by the scikit-learn maintainers. www.probabl.ai

Mapped capabilities

3 capabilities

  • Evaluation and validation methodology

    Whether the evaluation scheme behind a reported result is stated and appropriate to the task.

  • Metric selection and reporting

    Naming the metric a claimed improvement is measured on.

  • Guarding against unsupported claims

    Not asserting a model is better than the evidence in the run supports.

03

Explainability and traceability

Answering the why, not just the what: the site's stated failure mode is shipping a result you cannot explain, where traceability matters the moment a decision must be defended.

Pricing, deployment options (SaaS, private cloud, on-prem), security review, and a working demo on your data. www.probabl.ai

Mapped capabilities

3 capabilities

  • Decision rationale

    Explaining why the agent chose a given modeling step.

  • Path from data to result

    Tracing a reported outcome back through the pipeline steps that produced it.

  • Preserving user options

    Surfacing choices to the user instead of silently deciding on their behalf.

Illustrative example

Input
Pick the churn problem, improve the model, and tell me what you did.
Expected behavior
Skore reports the improved result together with the validation scheme and metric it was measured on, and attributes the gain to at least one specific pipeline change rather than returning only a final number.

04

Reproducibility and lineage

Directly targets the stated problem that hyperparameters, dataset versions, and model lineage live in scattered places, so last week's experiment cannot be rerun or compared with confidence.

Watch me improve a real ML model, live. No install, no account. www.probabl.ai

Mapped capabilities

4 capabilities

  • Hyperparameter capture

    Recording the parameters a run actually used.

  • Dataset versioning

    Binding a result to the version of the data it was produced from.

  • Model lineage

    Relating a model to the runs and changes it descends from.

  • Rerun and comparison

    Re-executing a prior experiment and comparing it to the current one.

Illustrative example

Input
Rerun the churn experiment I ran last week and tell me whether the improvement still holds on the current data.
Expected behavior
Skore locates the recorded run, restates its hyperparameters and dataset version, re-executes it, and reports a prior-versus-current comparison. It does not silently train a new, unlinked model and present that score as the answer.

05

Portability and deployment

The any AI provider, any ML framework, any compute claim, plus the delivery matrix of open-source `pip install skore`, SaaS, private cloud, and on-prem.

Any AI provider, any ML framework, any compute. www.probabl.ai

Mapped capabilities

4 capabilities

  • AI provider independence

    Behavior consistency when the underlying model provider changes.

  • ML framework coverage

    Operating over pipelines built in different ML frameworks.

  • Deployment mode fit

    Correctly describing what SaaS, private cloud, and on-prem each imply for a prospective team.

  • Open-source entry path

    Getting a user running from `pip install skore` without an account.

06

Trust, support, and procurement surface

The non-product surfaces a buyer touches: legal documents, privacy and DPO contacts, vulnerability disclosure, and the three declared routes in for sales, support, and open source.

Open source first. pip install skore and you’re up. www.probabl.ai

Mapped capabilities

4 capabilities

  • Legal document routing

    Pointing to the current Privacy Policy and Skore EULA versions and their review dates.

  • Procurement and security requests

    Directing DPA, security questionnaire, and SOC 2 report requests to the stated channel.

  • Support versus sales versus community

    Routing a user to GitHub issues, a demo booking, or Discord as their situation warrants.

  • Commitment accuracy

    Representing stated commitments such as a one business day reply without overstating them.

Coverage is mapped from Probabl's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Probabl test?+

The coverage map is generated from Probabl's own public product surface (AI data science agent / ML methodology platform): 6 scoring areas — Agentic modeling workflow, Scientific rigor and methodology, and Explainability and traceability, and more — spanning 22 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Probabl evals scored?+

Every case generated for Probabl — across Agentic modeling workflow and Scientific rigor and methodology and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Probabl library include?+

The full Probabl library is built on request. The coverage map spans 6 areas and 22 capabilities (for example, Problem loading and task framing and Pipeline drafting under Agentic modeling workflow); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Probabl or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Probabl areas and set them up in a Corsac workspace, where you can run every test case against Probabl or your own agent with your own data.