Pricing
Start Free. Scale When You're Ready.
$99 buys Pro: two monitors, 10 pooled runs a month, and Prompt CI on every deploy — 150 of your own prompts, 2 candidates, 5 judges rotated from a fixed pool. Model inference is included.
Free
one sponsored comparison
One sponsored comparison on your own prompts.
Start Sponsored Comparison- 1 sponsored comparison, judged and scored
- Standard & Advanced models
- Directional result — 5 judges on a handful of examples
- Reports guaranteed available for 7 days
- 1 seat
- Standing monitors
- Prompt CI
- API & MCP access
- Premium & Frontier models
Pro
per organization · unlimited seats
Every prompt and model change checked against your real traffic.
Upgrade to Pro- Prompt CI — deploy webhook and GitHub Action
- Up to 2 monitors · 10 organization-pooled runs/mo
- Weekly, plus catalog, price, drift and deploy triggers
- 150 of your own prompts, 2 candidates per run
- 5 judges rotated from a frozen pool of 8
- Model inference included within hosted-generation limits
- Verdicts to Slack, signed webhooks, and your tracing tool
- Premium models as candidates
- Unlimited seats
- API, MCP, and production traffic ingest
- Custom eval criteria
- Reports guaranteed available for 90 days
- Premium judge pool
Team
per organization · unlimited seats
Premium judges and the volume to gate every deploy.
Upgrade to Team- Everything in Pro
- Premium judge pool — the strongest evidence we'll seat
- Up to 6 monitors · 60 organization-pooled runs/mo
- Prompt-as-variable runs
- Reports guaranteed available for 180 days
- Choose your own judges
- Bring your own provider keys
- Frontier models as candidates
Enterprise
10 monitors · 100 pooled runs default
Frontier models, your own judges, and your own provider keys.
Start with a Free Managed TrialTalk to sales →- Everything in Team
- Frontier models as candidates
- Choose your own judges
- Bring your own provider keys — judges and candidates both
- 10 monitors and 100 pooled runs by default; contract may override
- Reports guaranteed available for 365 days
- Audit logs
- Dedicated support & SLA
- Free managed trial included
Need more? A run pack is $99/month and adds one monitor plus 10 pooled runs, up to four self-serve. That is the only add-on — inference is included and there is no per-run overage to opt into.
All paid plans billed monthly. No contracts — cancel anytime.
What a run includes
You buy monitors, runs, and triggers. PeerLM pays the model providers. There is no second, metered vocabulary to reconcile against the sticker price.
150 of your prompts
Each cycle samples your own production traffic — not a synthetic benchmark. Observed control uses the captured response; replayed control generates both sides. PeerLM rejects a run it cannot execute at the size you asked for rather than quietly shrinking it.
5 judges, frozen pool of 8
PeerLM selects the judges. A model from the vendor under test can't judge it. Five judges rotate from a fixed pool so consecutive cycles compare like with like.
Inference included
Generations and judging are bundled. We show provider list price as a range so you can see what the cycle costs us — not what you owe on top of the plan.
Runs are the ceiling
Runs are pooled across every Monitor in the organization. A run either happens or it doesn't. Unused runs do not roll over.
The premium judge pool is part of Team, not an add-on. Need more? A run pack is $99/month and adds one monitor plus 10 pooled runs, up to four self-serve. That is the only add-on — inference is included and there is no per-run overage to opt into.
Monitors — standing comparison on production traffic
Connect production traffic to a Monitor and continuously compare candidates for quality and cost. Standing monitors start on Pro. Free includes one sponsored comparison to try it out.
Sources
OpenTelemetry, Langfuse, Braintrust, LangSmith, Helicone, Cloudflare AI Gateway, direct API, and file upload.
Triggers
Weekly heartbeat, plus catalog, price, drift, and a deploy trigger on every paid plan — deploy Runs bypass the daily debounce.
Living verdict
Switch, route, or hold — with quality retained and projected savings evaluated against the decision contract. Evidence strength describes the support behind the result; it does not authorize a switch.
Routing export
Export recommended model routing as JSON, YAML, LiteLLM proxy config, Portkey gateway config, or a code snippet.