judgeLLM-judge assertions on every plan→

Know what your prompt change broke.

Veval turns real production traces into a regression baseline for your agents. When you change a prompt or swap a model, it shows you which steps behave differently, so you find out before your users do.

Free plan · No credit card · Python, Node.js & C# SDKs

code-review-agent / suite: main / run #184

scenario

sql-injection-concat

baseline

snapshot · pinned Apr 12

candidate

prompt v14

StepBaselineCandidate
classify-languagecsharpcsharp✓
check-security1 finding · High0 findings✕
check-style0 findings0 findings✓
check-logic1 finding · Medium1 finding · Medium✓
summarizeseverity: Highseverity: Medium✕
escalatecalledskipped✕
Judge

“Flags the SQL injection in the concatenated query.” fail · 0.12

The candidate review never mentions the unparameterized id and downgrades overall severity to Medium.

00The problem

Your checks all passed and the agent still got worse.

A one-line prompt edit can quietly change how an agent behaves three steps later. Your unit tests stay green because they test your code and say nothing about the model’s judgment.

Usually you find out from a support ticket. With Veval the regression shows up in the pull request.

prompts/check_security.txt+1 −1
  You are a senior security reviewer.- Report every vulnerability, however minor.+ Be concise. Only report issues that matter.  Respond with a JSON array of findings.
✓unit-tests214 passed
✓lintno issues
✓type-checkno errors
Merged into main · deployed

Two days later · #support

“Your reviewer approved a PR with raw SQL string concatenation. It used to catch that.”

01How it works

How a production trace becomes a regression test.

  1. 01

    Capture

    Wrap your agent with the SDK and each production run is recorded as a trace, one step at a time.

    14:02:11review · a91f3.2s
    14:02:09review · 7c0e2.8s
    14:01:58review · d34b4.1s
    14:01:40review · 02aa2.9s
  2. 02

    Pin

    When you find a run that went right, pin it. It becomes a snapshot your agent is held to from then on.

    trace a91fpinned

    scenario sql-injection-concat

    steps 6 · judge 1 rule

    retention never expires

  3. 03

    Detect

    Run the suite in CI. Veval replays each snapshot against your change and flags any step that behaves differently.

    ✕veval / regression suite

    2 regressed · 5 unchanged

    check-security: 1 finding → 0

Instrumenting takes a few lines and no new framework.

Your agent code, model provider and prompts stay where they are. The SDK records what your agent already does.

  • Python, Node.js and C#
  • Any model provider
  • Unlimited seats on every plan
Read the quickstart
1import veval2 3veval.init(api_key="...")4 5@veval.trace6def run_agent(prompt: str) -> str:7    ...

02Pricing

Free to start, with paid plans as your usage grows.

LLM-judge assertions are included on every plan. If you bring your own Anthropic or OpenAI key, you can judge with any model and it won’t use credits.

Free

For trying Veval on one agent.

$0/mo

Traces / mo
1k
Test runs / mo
50
Retention
7 days
Judge credits / mo
BYOK
Pinned snapshots
3
Start free

Starter

Popular

For a team shipping its first agent.

$29/mo

Traces / mo
10k
Test runs / mo
1k
Retention
14 days
Judge credits / mo
$5
Pinned snapshots
10
Get started

Growth

For agents with real traffic.

$99/mo

Traces / mo
100k
Test runs / mo
10k
Retention
30 days
Judge credits / mo
$25
Pinned snapshots
50
Get started

Pro

For several agents in production.

$299/mo

Traces / mo
250k
Test runs / mo
Unlimited
Retention
90 days
Judge credits / mo
$100
Pinned snapshots
200
Get started

→  Get started

Pin your first trace before your next prompt change.

Free plan · No credit card required