All evals
Venice

Eval directory

Evals for Venice

Eval coverage for Venice, mapped from its public product surface.

About Venice

Venice is a privacy-focused AI platform that provides chat, image, video, audio, code, and search generation through models from many leading providers, run in a private or anonymized mode. It positions itself as uncensored and private by default, encrypting content in the browser rather than storing it on Venice servers, with no account required to start. It sells Free, Pro ($18/mo), Pro Plus ($68/mo), and Max ($200/mo) plans plus an OpenAI-compatible API, and operates a crypto token ecosystem (VVV and DIEM) that can unlock Pro access and AI credits.

Industry

private, uncensored AI assistant and API platform

Website

venice.ai

Use the eval library for Venice

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Venice?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Model Selection & Routing

Helping a user choose among models from many providers (Claude, OpenAI, Google, DeepSeek, Mistral, Meta, Qwen, Grok, Kimi, Black Forest Labs, ElevenLabs, Runway, Kling, and others) and understand what changes when they switch.

Mapped capabilities

4 capabilities

  • Explaining differences between available models

    Capability and provider distinctions surfaced on the models and FAQ pages

  • Context lengths and extended context on paid tiers

    Per-model context limits; extended windows advertised for Pro

  • Privacy badges and TEE vs. E2EE distinctions

    What each badge asserts about how a given model runs

  • Switching models mid-conversation

    Stated effects on an in-progress chat

02

Multimodal Generation Workflows

The core creative surfaces the homepage advertises — text, image, video, audio, code, and search in one place — including the editing operations layered on image generation.

Create text, image, video, code, build agents, and more using fully private or anonymized models venice.ai

Mapped capabilities

4 capabilities

  • Text chat, reasoning, writing, and code generation

    Primary chat surface and code use cases

  • Image generation plus upscale, background removal, and variants

    Pro 'image superpowers' operations

  • Video generation from text or image prompts

    Access to Sora, Kling, Runway, Veo and similar models

  • Voice mode, music, and web search

    Enabling voice mode; how web search is described to work

03

Privacy, Encryption & Content Policy

Venice's central claim: prompts, responses, images, and uploads are encrypted in the browser and never stored on Venice servers, with no account required, alongside its uncensored positioning.

All your content - prompts, responses, images, document uploads - is encrypted in your local browser venice.ai

Mapped capabilities

4 capabilities

  • Browser-side encryption and no-server-storage claims

    Encryption, proxy circulation, and what Venice states it does not retain

  • No-account access and what it costs the user

    Tradeoffs of using Venice without signing up

  • Memory, history, and cross-device or cross-browser limits

    Where history lives; importing memories from another assistant

  • Uncensored positioning and its stated boundaries

    How 'uncensored and private by default' is described to users

Illustrative example

Input
Is my chat history stored on your servers? I never made an account, and I lost everything when I cleared my browser.
Expected behavior
Confirms prompts and responses are encrypted in the browser and not stored on Venice servers, so history lives locally in that browser and clearing site data loses it, and points to encrypted backup and restore as the way to preserve it.

04

Plans, Credits & Billing

The Free / Pro ($18) / Pro Plus ($68) / Max ($200) ladder, monthly vs. yearly pricing, and the credit system that governs video, music, frontier models, and API usage.

Text, image, video, audio, code, and search in one place, all private or anonymous. venice.ai

Mapped capabilities

4 capabilities

  • Tier limits on text and image prompts

    Free daily caps vs. Pro unlimited text and 1,000 images/day

  • Monthly credit allotments per tier

    100 / 7,500 / 22,500 credits and what credits are spent on

  • Credit banking and rollover windows

    2-month banking on Pro Plus, 3-month on Max

  • Monthly vs. yearly billing and upgrade paths

    Advertised yearly savings and tier-to-tier upgrades

Illustrative example

Input
I'm on Pro at $18 a month. How many credits do I get each month, and do unused credits roll over to the next month?
Expected behavior
States Pro includes 100 credits per month and that credit banking is a Plus and Max feature — two-month roll-forward on Pro Plus, three-month on Max — so Pro credits do not roll over. Names the upgrade tiers without inventing other rollover terms.

05

Developer API Enablement

The OpenAI-compatible API sold alongside the consumer app, where entitlements and rate limits are tied to subscription tier and credit balance.

Venice offers an OpenAI-compatible API for developers. venice.ai

Mapped capabilities

4 capabilities

  • OpenAI compatibility and migration expectations

    What 'OpenAI-compatible' does and does not imply for a port

  • API access by tier and rate limits

    Free-tier API for Pro; higher limits on Plus and Max

  • Credit-based access to premium models via API

    How credits meter API calls to frontier models

  • Routing developers to API docs and status page

    Docs, changelog, and status surfaces linked from the site

06

Token Ecosystem (VVV & DIEM)

The crypto layer that converts token holdings into product access: VVV as an ERC-20 on Base with staking and burns, and DIEM as a daily AI-credit instrument minted from staked VVV.

Mapped capabilities

4 capabilities

  • Staking VVV to unlock Pro access

    The stated 100 VVV staking threshold and resulting entitlements

  • DIEM minting and daily credit mechanics

    1 DIEM = $1 of daily AI credit; minting restricted to stakers

  • Tokenomics, supply, and the monthly revenue burn

    Launch supply, buy-and-burn program starting Nov 2025

  • Contract details and verification questions

    Contract address, staking contract control, and mint-authority FAQs

Coverage is mapped from Venice's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Venice test?+

The coverage map is generated from Venice's own public product surface (private, uncensored AI assistant and API platform): 6 scoring areas — Model Selection & Routing, Multimodal Generation Workflows, and Privacy, Encryption & Content Policy, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Venice evals scored?+

Every case generated for Venice — across Model Selection & Routing and Multimodal Generation Workflows and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Venice library include?+

The full Venice library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Explaining differences between available models and Context lengths and extended context on paid tiers under Model Selection & Routing); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Venice or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Venice areas and set them up in a Corsac workspace, where you can run every test case against Venice or your own agent with your own data.