PeerLM logoPeerLM
All Comparisons

Anthropic: Claude Sonnet 4.6 vs Anthropic: Claude Sonnet 4.5: Coding Performance with 10 Evaluators

We analyze the Coding Performance with 10 Evaluators to see how Anthropic: Claude Sonnet 4.6 vs Anthropic: Claude Sonnet 4.5 stack up against each other.

Anthropic: Claude Sonnet 4.6

4.3

preference score

vs

Anthropic: Claude Sonnet 4.5

5.7

preference score

Judges ranked the responses in this Run against each other; the rank is mapped onto a 0–10 scale. It shows which response was preferred, not how good either one is — and it is not a percentage, a pass rate, or a check that the output was correct.

Sample size for this comparison was not recorded. Treat it as directional.

Evidence clarification: this article predates recorded sample provenance. Treat its conclusions as claims about the displayed examples; they do not establish general model superiority, verified correctness, or production suitability.

Key Findings

Top RankAnthropic: Claude Sonnet 4.5

Secured the #1 spot in coding performance with a 5.68 overall score.

Instruction FollowingAnthropic: Claude Sonnet 4.5

Demonstrated superior consistency in adhering to complex coding prompts.

Cost EfficiencyAnthropic: Claude Sonnet 4.5

Maintains a lower total cost per response compared to the 4.6 iteration.

Specifications

SpecAnthropic: Claude Sonnet 4.6Anthropic: Claude Sonnet 4.5
Provideranthropicanthropic
Context Length1.0M1.0M
Input Price (per 1M tokens)$3.00$3.00
Output Price (per 1M tokens)$15.00$15.00
Max Output Tokens128,00064,000
Tierfrontierfrontier

Our Verdict

Anthropic: Claude Sonnet 4.5 is the clear winner in this coding-focused evaluation, outperforming the 4.6 iteration in both accuracy and instruction following. Developers looking for the most reliable coding assistant should prioritize the 4.5 model based on these PeerLM benchmark results.

Overview

In this evaluation, we take a deep dive into the performance of two prominent models from Anthropic: Claude Sonnet 4.6 and Claude Sonnet 4.5. This comparative analysis focuses specifically on Coding Performance with 10 Evaluators, utilizing PeerLM's rigorous testing framework to determine how these iterations handle real-world programming tasks and complex instructions.

Benchmark Results

When looking at the Coding Performance with 10 Evaluators, the rankings highlight a distinct leader in the current evaluation cycle. PeerLM's comparative approach allows us to look past simple metrics and focus on how these models perform when ranked by human-aligned evaluators.

ModelRankOverall ScoreAccuracyInstruction Following
Anthropic: Claude Sonnet 4.515.685.685.68
Anthropic: Claude Sonnet 4.624.324.324.32

Criteria Breakdown

The evaluation centered on two critical pillars for any coding assistant: Accuracy and Instruction Following. In coding, accuracy determines the functional correctness of the generated logic, while instruction following ensures that the model adheres to specific architectural constraints, style guides, or framework requirements provided by the user.

The current results show a score spread of 1.36 between the two models. Anthropic: Claude Sonnet 4.5 demonstrates a more refined ability to meet the expectations of our 10 evaluators, maintaining a higher level of consistency across both criteria compared to the 4.6 iteration.

Cost & Latency

Understanding the economic and temporal cost of model deployment is essential for developers. Below is the breakdown of the resource usage for these models during the evaluation.

ModelAvg Latency (ms)Total Cost (USD)Avg Completion Tokens
Anthropic: Claude Sonnet 4.519530.014019186
Anthropic: Claude Sonnet 4.600.014196189

While Anthropic: Claude Sonnet 4.5 shows a measurable latency of 1953ms, it maintains a slightly lower total cost per response compared to the 4.6 model. These metrics are vital for teams scaling their internal coding agents or building production-grade software.

Use Cases

Given the results of the Coding Performance with 10 Evaluators, Anthropic: Claude Sonnet 4.5 is currently the preferred choice for complex coding tasks where precision is paramount. Its higher ranking suggests it is better suited for:

  • Writing complex boilerplate code in new frameworks.
  • Refactoring legacy systems where instruction adherence is critical.
  • Debugging logic errors that require multi-step reasoning.

Anthropic: Claude Sonnet 4.6, while trailing in this specific evaluation, remains a robust option for general-purpose tasks and iterative drafting where rapid prototyping is prioritized.

Verdict

Based on our comparative evaluation, Anthropic: Claude Sonnet 4.5 outperforms the 4.6 version in the specific context of coding tasks. The 1.36 score gap suggests that for developers requiring the highest fidelity in code generation and adherence to technical instructions, the 4.5 model currently provides a more reliable output.

Backed by real data

View the Full Evaluation Report

See every response, score, and evaluator judgment behind this comparison. All data from PeerLM's blind evaluation pipeline.

View Report

Run your own Monitor

Compare Anthropic: Claude Sonnet 4.6 and Anthropic: Claude Sonnet 4.5 on sampled production prompts, with frozen criteria and inspectable evidence.

Start a Monitor

Get a free managed report

We'll run a full evaluation with your real prompts and deliver a detailed recommendation. Free for qualified teams.

Request Report

Methodology

Evaluated using PeerLM's blind evaluation pipeline with 4 responses per model across 2 criteria.