All evals
MA

Eval directory

Evals for Megaton AI

Eval coverage for Megaton AI, mapped from its public product surface.

About Megaton AI

Megaton is a web app and directory for AI video generation, combining model leaderboards, benchmark scores, audited per-second pricing pages, and AI video news coverage. It also sells credit-based video editing tools including background removal (green screen), AI video masking, and object/watermark removal. Access is pay-as-you-go via credit packs rather than a subscription, and 'Try' links to third-party models may be affiliate links.

Industry

AI video model marketplace, benchmarks and editing tools

Website

megaton.ai

Use the eval library for Megaton AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Megaton AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Model leaderboard and rankings

The ranked directory of video models surfaced across the site — top-model rail, full leaderboard, and the overall scores attached to each model.

We benchmark dozens of models on quality, speed, and value—then offer only the best. megaton.ai

Mapped capabilities

4 capabilities

  • Rank and score consistency across surfaces

    Same model's rank and overall score agree between the homepage rail, leaderboard, and its pricing page.

  • Tie and ordering handling

    Equal scores that share a rank (e.g. two models at #3) are presented without implying a false ordering or a skipped position.

  • Curation scope and omissions

    Explains that unreleased or below-standard models are absent rather than implying the list is exhaustive.

  • Quality-versus-cost framing

    Superlatives like Quality Leader, Best Value, and Fastest map to the dimension they claim, not to overall rank alone.

02

Per-second pricing pages

Audited per-second rate pages for individual models, including normalized clip estimates, comparison cards, and pricing FAQs.

Mapped capabilities

4 capabilities

  • Clip-length arithmetic

    5s, 10s, 30s, 60s, and custom-length estimates equal the published per-second rate times duration, with consistent rounding.

  • Rate-to-headline agreement

    Page title, displayed price per second, per-minute figure, and FAQ answers all cite the same audited rate.

  • Comparison card correctness

    Adjacent model cards show each peer's own rate and score, not the host model's, and link to the right pricing page.

  • Scope of what the price covers

    States access channel, published resolution, and audio inclusion without inventing resolution markups the source does not publish.

Illustrative example

Input
What would a 30-second Kling 3 Pro clip cost, and how does that compare to Seedance 2.0 at the same length?
Expected behavior
Answers $5.04 for Kling 3 Pro at $0.168 per second and $4.54 for Seedance 2.0 at $0.1512 per second, noting Seedance 2.0 is cheaper and rates are a July 22, 2026 snapshot.

03

Credit packs and purchase economics

Pay-as-you-go credit packs (Starter, Builder, Scale) with per-credit rates, savings claims, and usage estimates.

Mapped capabilities

4 capabilities

  • Per-credit and savings math

    Per-credit price and stated savings percentages follow from pack dollar amount and credit count.

  • Usage estimate consistency

    Approximate counts of background removals, masking minutes, and generations track the stated per-operation credit costs.

  • Pack recommendation for a stated workload

    Given a described volume, recommends a pack and shows the arithmetic rather than defaulting to the largest.

  • No-subscription and expiry terms

    Represents credits-never-expire, no-subscription, and all-features-included terms as published.

Illustrative example

Input
I expect about 400 background removals a month. Which credit pack should I buy?
Expected behavior
Recommends the $25 Builder pack, whose 3,000 credits cover roughly 600 background removals, and notes the Starter pack covers only about 100. Mentions credits never expire and there is no subscription.

04

Video editing tools

The credit-metered editing capabilities sold in the app: green screen background removal, AI video masking, and paint-out object/watermark removal.

Buy credits once, use them for any operation. Prices vary by model and task. megaton.ai

Mapped capabilities

4 capabilities

  • Tool-to-task routing

    Directs a described editing goal to green screen, masking, or paint-out per each tool's stated purpose.

  • Per-operation cost quoting

    Quotes the published from-price bands and flags that cost varies by model and task.

  • Per-minute versus per-operation metering

    Distinguishes masking priced per minute from removal operations priced per run when estimating a job.

  • Capability boundaries

    Does not attribute editing features to Megaton beyond background removal, masking, paint-out, and generation.

05

AI video news and model coverage

Dated editorial output across Regulation, Culture, and model evaluation categories, including long-form model write-ups.

Mapped capabilities

4 capabilities

  • Article attribution and dating

    Headline, category, byline, and publication date are reported together and accurately.

  • Category routing

    A topical query returns coverage from the matching category rather than the newest item overall.

  • Model evaluation summarization

    Summarizes an evaluation's claims about the model without inflating them into benchmark scores.

  • Recency boundaries

    Declines to assert coverage of events after the latest published article date.

06

Sourcing, disclosure, and freshness

The trust layer around the data: affiliate disclosure on Try links, pricing source type, confidence, and last-checked dates.

“Try” links may be affiliate links — Megaton may earn a commission at no extra cost to you. megaton.ai

Mapped capabilities

4 capabilities

  • Affiliate disclosure on Try links

    Surfaces the affiliate-commission disclosure whenever a Try link to a third-party model is recommended.

  • Source provenance labeling

    Reports source type (official API versus third-party API), access channel, and confidence as published per model.

  • Staleness handling

    Cites the last-checked date and treats rates as a snapshot rather than a live quote.

  • Unpublished-figure refusal

    Declines to supply resolution markups, latency numbers, or rates the pricing snapshot does not record.

Coverage is mapped from Megaton AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Megaton AI test?+

The coverage map is generated from Megaton AI's own public product surface (AI video model marketplace, benchmarks and editing tools): 6 scoring areas — Model leaderboard and rankings, Per-second pricing pages, and Credit packs and purchase economics, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Megaton AI evals scored?+

Every case generated for Megaton AI — across Model leaderboard and rankings and Per-second pricing pages and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Megaton AI library include?+

The full Megaton AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Rank and score consistency across surfaces and Tie and ordering handling under Model leaderboard and rankings); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Megaton AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Megaton AI areas and set them up in a Corsac workspace, where you can run every test case against Megaton AI or your own agent with your own data.