PeerLM logoPeerLM

Everything You Need to Know

How blind evaluation works, what Monitors deliver, what a plan includes, and why teams switch from manual testing.

What is blind comparative ranking?

Each challenger is compared pairwise with your control in both response positions, by a cross-provider panel that excludes the vendors under test. Running both positions catches judges that favour whichever answer came first, and the exclusion stops a vendor grading its own work. These controls reduce identifiable sources of evaluator bias without claiming to eliminate them.

How is PeerLM different from testing models manually?

Manual testing often means one person, one prompt, and a decision that is hard to reproduce. PeerLM runs blind evaluations across many scenarios, personas, and criteria, preserving the prompts, outputs, judge results, and parameters used. You get reviewable evidence instead of anecdotes, delivered to Slack with the receipt attached.

What models can I evaluate?

200+ models across OpenAI, Anthropic, Google, Meta, Mistral, and more — via OpenRouter and Groq integrations. Evaluate any combination against each other. Model capabilities sync automatically so you always have access to the latest releases.

What do I pay for?

Monitors, runs, and triggers. Model inference is included — PeerLM pays the providers, and we do not pass a token bill through to you. Free includes one sponsored comparison with 5 judges on a handful of examples; it is directional, not a standing Monitor. $99/month buys Pro: 2 monitors, 10 organization-pooled runs, Prompt CI on every deploy, weekly plus catalog, price and drift triggers, 150 of your own prompts, 2 candidates, and 5 judges rotated from a frozen pool of 8. Team is $499/month and adds the premium judge pool, 6 monitors, and 60 pooled runs. A run pack is $99/month for one more monitor and 10 more runs — the only add-on. Unused runs do not roll over.

What happens when I subscribe to a plan?

Your workspace is activated immediately. All evaluations and data are retained — nothing resets. Pro unlocks standing monitors, Prompt CI, Premium candidates, custom criteria, outbound verdicts, and API & MCP access. Team adds the premium judge pool. Start running comparisons right away. Reach out to sales@peerlm.com if you have any questions.

What are Monitors?

A Monitor connects PeerLM to a production use case and accumulates comparison Runs against live traffic. Each Run samples real prompts, replays candidates on the same inputs, and judges quality retained vs your control alongside cost. The Monitor rolls those Runs into a living switch / route / hold verdict. Standing monitors start on Pro. Free includes one sponsored comparison on your own prompts.

What's the difference between observed and replayed control?

Observed control is the production default: PeerLM compares each challenger with the exact captured production response and makes no incumbent provider call. It requires at least 90% usable output coverage plus a resolved incumbent and system-prompt boundary; otherwise the Run is blocked. Replayed control explicitly generates the frozen incumbent and every challenger as they behave today. Prompt CI replays the old and new prompt on the same model. PeerLM never changes control modes automatically.

What production traffic can PeerLM compare today?

Production Monitor comparisons currently support replayable single-turn request-and-response records with one resolved incumbent model and a known system-prompt boundary. Flattened multi-turn conversations, tool calls, and agent traces are not treated as replayable, and PeerLM refuses a Run rather than scoring traffic it cannot faithfully replay. Structured agent-step support is in development.

What is a living verdict?

The living verdict is the Monitor's recommendation — switch, route, or hold — refreshed by every comparison Run. Quality retained, latency, and projected savings are evaluated against the Monitor's frozen decision contract. Evidence strength is a separate 0–100 description of sample size, judge agreement, and run cleanliness; it contains no savings or quality terms and does not authorize a switch. Results break down by prompt cluster, and routing recommendations can be exported as JSON, YAML, LiteLLM config, Portkey config, or a code snippet. Receipts live on the Run; the verdict lives on the Monitor.

What log sources does PeerLM connect to?

OpenTelemetry is the recommended path — point any OTLP exporter at PeerLM and it works with Vercel AI SDK, OpenLLMetry, and OpenInference. PeerLM also pulls from Langfuse, Braintrust, LangSmith, Helicone, and Cloudflare AI Gateway, and accepts file upload or direct API ingest. Connect Sources on each Monitor, then start a comparison Run from that traffic.

What triggers an Auto-Run?

On every paid plan, five events can start another comparison Run: the weekly heartbeat, a model release (catalog), a price change, quality drift, and a deploy. Deploy triggers bypass the daily debounce so each ship is its own run. Runs are pooled across the organization (10 on Pro, 60 on Team; Enterprise defaults to 100 and may override in the contract). Configure on the Monitor → Schedule.

Who picks the judges?

PeerLM selects the judges. A model from the vendor under test can't judge it. Every cycle seats 5 judges rotated from a frozen pool of 8, so consecutive runs compare like with like — the panel always runs in full. The sponsored comparison also seats 5 judges but remains directional because its sample is small. Team includes the premium judge pool, which you assign per Monitor when a decision needs the strongest evidence we'll seat. Choosing your own judges is available on Enterprise.

What is Prompt CI?

On every paid plan, a deploy webhook and GitHub Action validate every system-prompt change against real production traffic before it ships. The control is the old prompt on the same model — not a cheaper candidate. Deploy runs bypass the daily debounce and count against the organization's pooled run allowance.

What are system prompts?

System prompts (personas) let you test models in context — as a customer support agent, code reviewer, creative writer, etc. Each system prompt defines the role and instructions given to the model. PeerLM breaks down results by system prompt so you see which model wins for each use case.

What's the difference between JSON output and text output?

JSON mode returns structured arrays (e.g., '3 one-liner jokes') where each item is scored independently — giving granular data. Text mode evaluates free-form prose as a whole. JSON mode is recommended for comparative evaluations as it produces more reliable rankings.

Can I share results with my team?

Yes. All paid plans include shareable reports — a public link anyone can view without an account. Reports include rankings, per-persona breakdowns, and response samples. Export as CSV or JSON for further analysis.

How does response caching work?

PeerLM hashes the combination of model, persona, and topic content. Matching cached responses are reused instead of regenerated. Edit any part of the prompt and the cache auto-invalidates. Caching makes runs faster; it does not change your run allowance.

What is deterministic mode?

PeerLM attempts temperature 0 and fixed seeds where each model supports them. It checks the capability registry first and omits unsupported parameters. The rendered prompt and parameters actually used are logged for audit purposes, improving repeatability without claiming every provider is deterministic.

Do you store my API keys?

Model calls use PeerLM-managed keys by default — no setup needed. On Enterprise, Bring Your Own Key covers both the judge panel and the models under test. Keys are encrypted with AES-256-GCM at rest under a dedicated key, and only the last four characters are ever readable.

What is the PeerLM MCP server?

The PeerLM MCP (Model Context Protocol) server lets you run evaluations directly from Claude Desktop, Cursor, or any MCP-compatible client. Connect traffic, trigger comparison Runs, and read verdicts — all without leaving your IDE. Available on Pro, Team, and Enterprise. Set it up with: npx -y @peerlm/mcp

Does PeerLM have an API?

Yes. The REST API lets you manage Monitors, ingest production traffic, start comparison Runs, validate a prompt change through Prompt CI, and retrieve results programmatically. Available on Pro, Team, and Enterprise. Generate an API key from Settings > API Keys.

Can I cancel anytime?

Yes. Month-to-month, no long-term commitment. Cancel or downgrade anytime from billing settings. You keep access through the end of your billing period.

Still have questions?

Contact support