Know what your prompt change broke.
Veval turns real production traces into a regression baseline for your agents. When you change a prompt or swap a model, it shows you which steps behave differently, so you find out before your users do.
Free plan · No credit card · Python, Node.js & C# SDKs
code-review-agent / suite: main / run #184
5 passedscenario
sql-injection-concat
baseline
snapshot · pinned Apr 12
candidate
prompt v14
| Step | Baseline | Candidate | |
|---|---|---|---|
| classify-language | csharp | csharp | ✓ |
| check-security | 1 finding · High | 0 findings | ✕ |
| check-style | 0 findings | 0 findings | ✓ |
| check-logic | 1 finding · Medium | 1 finding · Medium | ✓ |
| summarize | severity: High | severity: Medium | ✕ |
| escalate | called | skipped | ✕ |
“Flags the SQL injection in the concatenated query.” fail · 0.12
The candidate review never mentions the unparameterized id and downgrades overall severity to Medium.
00The problem
Your checks all passed and the agent still got worse.
A one-line prompt edit can quietly change how an agent behaves three steps later. Your unit tests stay green because they test your code and say nothing about the model’s judgment.
Usually you find out from a support ticket. With Veval the regression shows up in the pull request.
You are a senior security reviewer.- Report every vulnerability, however minor.+ Be concise. Only report issues that matter. Respond with a JSON array of findings.Two days later · #support
“Your reviewer approved a PR with raw SQL string concatenation. It used to catch that.”
01How it works
How a production trace becomes a regression test.
- 01
Capture
Wrap your agent with the SDK and each production run is recorded as a trace, one step at a time.
14:02:11review · a91f3.2s14:02:09review · 7c0e2.8s14:01:58review · d34b4.1s14:01:40review · 02aa2.9s - 02
Pin
When you find a run that went right, pin it. It becomes a snapshot your agent is held to from then on.
trace a91fpinnedscenario sql-injection-concat
steps 6 · judge 1 rule
retention never expires
- 03
Detect
Run the suite in CI. Veval replays each snapshot against your change and flags any step that behaves differently.
✕veval / regression suite2 regressed · 5 unchanged
check-security: 1 finding → 0
Instrumenting takes a few lines and no new framework.
Your agent code, model provider and prompts stay where they are. The SDK records what your agent already does.
- Python, Node.js and C#
- Any model provider
- Unlimited seats on every plan
1import veval2 3veval.init(api_key="...")4 5@veval.trace6def run_agent(prompt: str) -> str:7 ...02Pricing
Free to start, with paid plans as your usage grows.
LLM-judge assertions are included on every plan. If you bring your own Anthropic or OpenAI key, you can judge with any model and it won’t use credits.
Free
For trying Veval on one agent.
$0/mo
- Traces / mo
- 1k
- Test runs / mo
- 50
- Retention
- 7 days
- Judge credits / mo
- BYOK
- Pinned snapshots
- 3
Starter
PopularFor a team shipping its first agent.
$29/mo
- Traces / mo
- 10k
- Test runs / mo
- 1k
- Retention
- 14 days
- Judge credits / mo
- $5
- Pinned snapshots
- 10
Growth
For agents with real traffic.
$99/mo
- Traces / mo
- 100k
- Test runs / mo
- 10k
- Retention
- 30 days
- Judge credits / mo
- $25
- Pinned snapshots
- 50
Pro
For several agents in production.
$299/mo
- Traces / mo
- 250k
- Test runs / mo
- Unlimited
- Retention
- 90 days
- Judge credits / mo
- $100
- Pinned snapshots
- 200
→ Get started
Pin your first trace before your next prompt change.
Free plan · No credit card required