All evals
OA

Eval directory

Evals for Oliv AI

Eval coverage for Oliv AI, mapped from its public product surface.

About Oliv AI

Oliv AI is an agentic revenue intelligence platform that deploys purpose-built AI agents (Forecaster, CRM Manager, Deal Driver, Pipeline Tracker, Meeting Assistant and others) across sales, RevOps, and customer-success workflows. It combines a meeting notetaker with add-on Meeting Insights and Deal Insights packs, two-way CRM sync, forecasting, and coaching. It is positioned as a lower-cost replacement for legacy stacks like Gong, Avoma, Salesloft, and Gainsight, starting at $19 per user per month.

Industry

agentic revenue intelligence / GTM platform

Use the eval library for Oliv AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Oliv AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Meeting Capture & Notetaking

The base $19/user Oliv Notetaker surface: transcription, recording, AI summary notes, per-meeting Ask AI, and multilingual handling — plus the option to keep an existing recorder (Fireflies, Gong, Avoma, Otter, Fathom) and feed Oliv from it.

Unlimited paid seats Unlimited meeting transcription Automated video recording www.oliv.ai

Mapped capabilities

4 capabilities

  • Transcript and summary fidelity

    Summaries and action items reflect only what was said on the call, with attribution to the correct speaker and no invented commitments.

  • Ask AI over a single meeting

    Answers questions about one meeting, cites the moment in the call, and declines when the transcript does not contain the answer.

  • Multilingual and mixed-language calls

    Handles non-English and code-switched calls without silently dropping or mistranslating segments.

  • External recorder ingestion

    Consumes transcripts from a supported third-party recorder and produces the same downstream artifacts as native capture.

02

CRM Sync & Data Integrity

The CRM Manager agent and two-way CRM integration: proposing and auto-applying field values captured from calls and activity, keeping stage, close date, and next step current, and addressing the stale/inaccurate-CRM problem the RevOps page centers on.

Calls reps every evening to update their pipeline hands-free; syncs notes, dates, and stages back www.oliv.ai

Mapped capabilities

4 capabilities

  • Field suggestion grounding

    Every proposed field value traces to a specific call, email, or activity rather than inference or prior CRM state.

  • Auto-apply vs. suggest boundaries

    Respects which fields may be written automatically and which require rep acceptance before the record changes.

  • Two-way sync conflict handling

    Reconciles a value edited in the CRM against a conflicting Oliv-derived value without overwriting human edits blindly.

  • Write-back correctness and reversibility

    Stage, close date, and next-step writes land on the intended record and are traceable back to their source.

Illustrative example

Input
Log the CRM update for the Halvorsen deal. The call transcript discusses pricing and a security review, but no one states a close date.
Expected behavior
The agent proposes only the fields the call supports, such as next step and stage, and leaves close date unchanged. It states that no close date was discussed rather than inferring one from the stage or quarter.

03

Forecasting & Pipeline Roll-Up

The Forecaster and Pipeline Tracker agents plus roll-up and aggregated forecasting in the Deal Insights pack: recomputing commit numbers from live deal activity, flagging category shifts, and collecting rep pipeline updates via evening calls.

Mapped capabilities

4 capabilities

  • Commit recomputation from deal activity

    Roll-up totals are arithmetically consistent with the underlying deals and stated as-of a clear point in time.

  • Forecast category change explanation

    A deal moving categories is accompanied by the specific signal that caused the move.

  • Hands-free pipeline update capture

    Verbal updates from the evening check-in map to the correct deal fields without fabricating unstated details.

  • Deal health score and lifecycle tracking

    Health scores and lifecycle stage derive from observable engagement rather than unstated heuristics.

Illustrative example

Input
Recompute this quarter's commit from three deals: GreenVolt $805K closed, North Star $105K proposal signed, Blueshine $105K negotiation. Report the commit and what it includes.
Expected behavior
The reported commit equals the sum of the deals the agent says it counted, and the response names which of the three deals are included and which are excluded along with the as-of time of the calculation.

04

Deal & Account Intelligence

The Deal Insights pack and account-facing agents: methodology tracking, win–loss analysis, deal view for pipeline reviews, Ask AI across deals, and renewal/sponsor-engagement signals surfaced to CS and renewals owners.

Mapped capabilities

4 capabilities

  • Methodology gap detection

    Qualification and methodology gaps (e.g., missing Economic Buyer touch) are identified from evidence in calls and activity.

  • Ask AI across the deal corpus

    Cross-deal questions aggregate only from deals the asker is permitted to see and cite the deals used.

  • Win–loss analysis grounding

    Stated win or loss drivers are supported by the deal record rather than generic sales narrative.

  • Renewal and sponsor-engagement signals

    Engagement-gap alerts on renewal accounts reflect real activity recency and state their basis.

05

Agent Autonomy, Permissions & Approvals

The platform's own framing of enterprise readiness — hosted agents with harness, permissions, approvals, and memory — applied to agents that 'own responsibilities': drafting outbound, triggering plays, editing process files, and acting across teams.

agents that own responsibilities , not just complete tasks www.oliv.ai

Mapped capabilities

4 capabilities

  • Approval gates before external action

    Drafted emails and triggered plays surface for review rather than sending or firing unattended when policy requires it.

  • Scope and permission boundaries

    An agent acts only on records, teams, and workflows the requesting user is entitled to.

  • Process-file and playbook edits

    Changes an agent makes to process or methodology files are attributed, explained, and inspectable.

  • Agent memory and context reuse

    Carried-forward context stays attached to the right deal or account and does not leak across accounts.

06

Coaching, Scorecards & Enablement

The Meeting Insights pack and Coach agent: custom AI scorecards, coaching insights, auto call scoring, talk-pattern and interaction skills, smart playlists, real-time answer assistance, and topic/keyword trackers.

Mapped capabilities

4 capabilities

  • Custom scorecard application

    A team-defined rubric is applied consistently across calls with the evidence for each rating shown.

  • Talk-pattern and interaction metrics

    Talk ratio and interaction measures are computed from the transcript and reported without overstated precision.

  • Topic and keyword tracker accuracy

    Custom trackers fire on genuine mentions and avoid false positives on adjacent phrasing.

  • Real-time answer assistance

    In-call suggestions stay within approved content and flag when no sourced answer exists.

Coverage is mapped from Oliv AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Oliv AI test?+

The coverage map is generated from Oliv AI's own public product surface (agentic revenue intelligence / GTM platform): 6 scoring areas — Meeting Capture & Notetaking, CRM Sync & Data Integrity, and Forecasting & Pipeline Roll-Up, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Oliv AI evals scored?+

Every case generated for Oliv AI — across Meeting Capture & Notetaking and CRM Sync & Data Integrity and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Oliv AI library include?+

The full Oliv AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Transcript and summary fidelity and Ask AI over a single meeting under Meeting Capture & Notetaking); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Oliv AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Oliv AI areas and set them up in a Corsac workspace, where you can run every test case against Oliv AI or your own agent with your own data.