All evals
DatologyAI

Eval directory

Evals for DatologyAI

Eval coverage for DatologyAI, mapped from its public product surface.

About DatologyAI

DatologyAI is a data curation platform (a "data refinery") that turns a customer's public, proprietary, and licensed data into high-quality training datasets for pre-training or mid-training models. Its pipeline has four stages — clean, curate, create (synthetic data), and compose — and runs inside the customer's own environment. It markets itself to AI model builders, with published research and customer case studies such as Arcee AI and Thomson Reuters.

Industry

AI training data curation platform

Use the eval library for DatologyAI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for DatologyAI?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Clean — corpus hygiene and decontamination

The first pipeline stage: ingesting heterogeneous customer sources and removing data that carries no training signal or contaminates evaluation, at petabyte scale.

We don't sell data. We don't source tokens. We make your data better. www.datologyai.com

Mapped capabilities

4 capabilities

  • Source ingestion and schema unification

    Connecting customer sources, unifying schema, repairing encoding across public, proprietary, and licensed inputs.

  • Heuristic filtering of degenerate samples

    Removing malformed, empty, or short documents and outliers such as extreme symbol-to-word ratios.

  • Benchmark decontamination

    N-gram leakage checks across all training sources so eval benchmarks are not contaminated.

  • Scale behavior of cleaning

    How cleaning is described as operating at petabyte scale rather than on sampled subsets.

Illustrative example

Input
We're pre-training on a web crawl plus our internal engineering docs. Does your platform handle test-set leakage, and where in the pipeline does that happen?
Expected behavior
Identifies decontamination as part of the Clean stage and describes it as an n-gram leakage check applied across all training sources, alongside ingestion and heuristic filtering. Does not name specific benchmarks or leakage rates that the published material never states.

02

Curate — shaping the training distribution

The second stage: understanding what is in the corpus and reshaping it by quality, task, and taxonomy so the retained data matches the model's intended job.

Datology is the data refinery for AI model builders. www.datologyai.com

Mapped capabilities

4 capabilities

  • Quality and taxonomy annotation

    Heuristic and learned classifiers that label corpus quality and topical taxonomy.

  • Redundancy reduction

    N-gram and embedding-based similarity rejection of near-duplicate samples.

  • Quality-based resampling

    Adjusting the sampling distribution to emphasize higher-quality data.

  • Task distribution matching

    Reshaping the corpus toward customer use cases using customer-supplied examples.

03

Create — synthetic data generation

The third stage: generating synthetic data to close coverage gaps that real data does not fill, as described in the BeyondWeb work on trillion-scale synthetic pretraining data.

Mapped capabilities

4 capabilities

  • Targeted document rephrasing

    BeyondWeb's rephrasing-based generation for diverse, relevant, information-dense synthetic pretraining data.

  • Task relevance and diversity coverage

    Filling coverage gaps without collapsing into low-diversity output.

  • Known synthetic-data failure modes

    How the approach is positioned against the failure modes attributed to naive synthetic generation.

  • Scaling evidence for synthetic data

    Claims about behavior at trillion-token scale relative to public pretraining datasets.

04

Compose — staged training data mixes

The fourth stage: mixing sources into staged datasets, including the distinction between mid-training on open models and pre-training a foundation model from scratch.

N-gram leakage check across all training sources www.datologyai.com

Mapped capabilities

4 capabilities

  • Source mixing across public, proprietary, and licensed data

    Composing an optimal mixture from the customer's three input classes.

  • Staged dataset construction

    Producing multi-stage datasets aligned to phases of a training run.

  • Mid-training vs. pre-training path selection

    When domain adaptation via mid-training on open models is the described path versus full pre-training.

  • Specialized-domain data placement

    The specialized-pretraining findings on repeating scarce domain data from pretraining onward.

05

Deployment model and data boundaries

How the platform is positioned operationally: running inside the customer's private cloud, operated by the customer's team, with Datology not selling or sourcing data.

all deployed on your private cloud www.datologyai.com

Mapped capabilities

3 capabilities

  • Private-cloud, in-environment execution

    The pipeline running in the customer's environment rather than as a hosted data service.

  • Customer-operated workflow

    The customer's own team running a frontier-level data process on the platform.

  • Data ownership boundary

    The stated position that Datology does not sell data or source tokens, only improves the customer's data.

06

Evidence, research, and case-study claims

The public proof surface: research posts and named customer case studies that a buyer would interrogate before trusting performance claims.

Mapped capabilities

4 capabilities

  • Customer case-study results

    Arcee AI and Thomson Reuters outcomes as published, including the figures actually stated.

  • Published research corpus

    Research updates spanning multilingual curation, VLM curation, embeddings, and evaluation work.

  • Compute-multiplier positioning

    The claim that better data quality raises signal per token and thereby stretches a fixed compute budget.

  • Claim attribution and scope limits

    Distinguishing what the published evidence covers from extrapolation to other domains or workloads.

Illustrative example

Input
What did the Thomson Reuters legal work actually show, and how expensive was it compared to pre-training from scratch?
Expected behavior
Reports the published figures — roughly +5% on LegalBench from mid-training against a proprietary legal corpus, at 100B mid-training tokens versus 15T pre-training, under 1% of the base pre-training budget. Attributes them to the case study and does not generalize to other domains.

Coverage is mapped from DatologyAI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for DatologyAI test?+

The coverage map is generated from DatologyAI's own public product surface (AI training data curation platform): 6 scoring areas — Clean — corpus hygiene and decontamination, Curate — shaping the training distribution, and Create — synthetic data generation, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the DatologyAI evals scored?+

Every case generated for DatologyAI — across Clean — corpus hygiene and decontamination and Curate — shaping the training distribution and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the DatologyAI library include?+

The full DatologyAI library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Source ingestion and schema unification and Heuristic filtering of degenerate samples under Clean — corpus hygiene and decontamination); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against DatologyAI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped DatologyAI areas and set them up in a Corsac workspace, where you can run every test case against DatologyAI or your own agent with your own data.