Overview
In this evaluation, we take a deep dive into the performance of two prominent models from Anthropic: Claude Sonnet 4.6 and Claude Sonnet 4.5. This comparative analysis focuses specifically on Coding Performance with 10 Evaluators, utilizing PeerLM's rigorous testing framework to determine how these iterations handle real-world programming tasks and complex instructions.
Benchmark Results
When looking at the Coding Performance with 10 Evaluators, the rankings highlight a distinct leader in the current evaluation cycle. PeerLM's comparative approach allows us to look past simple metrics and focus on how these models perform when ranked by human-aligned evaluators.
| Model | Rank | Overall Score | Accuracy | Instruction Following |
|---|---|---|---|---|
| Anthropic: Claude Sonnet 4.5 | 1 | 5.68 | 5.68 | 5.68 |
| Anthropic: Claude Sonnet 4.6 | 2 | 4.32 | 4.32 | 4.32 |
Criteria Breakdown
The evaluation centered on two critical pillars for any coding assistant: Accuracy and Instruction Following. In coding, accuracy determines the functional correctness of the generated logic, while instruction following ensures that the model adheres to specific architectural constraints, style guides, or framework requirements provided by the user.
The current results show a score spread of 1.36 between the two models. Anthropic: Claude Sonnet 4.5 demonstrates a more refined ability to meet the expectations of our 10 evaluators, maintaining a higher level of consistency across both criteria compared to the 4.6 iteration.
Cost & Latency
Understanding the economic and temporal cost of model deployment is essential for developers. Below is the breakdown of the resource usage for these models during the evaluation.
| Model | Avg Latency (ms) | Total Cost (USD) | Avg Completion Tokens |
|---|---|---|---|
| Anthropic: Claude Sonnet 4.5 | 1953 | 0.014019 | 186 |
| Anthropic: Claude Sonnet 4.6 | 0 | 0.014196 | 189 |
While Anthropic: Claude Sonnet 4.5 shows a measurable latency of 1953ms, it maintains a slightly lower total cost per response compared to the 4.6 model. These metrics are vital for teams scaling their internal coding agents or building production-grade software.
Use Cases
Given the results of the Coding Performance with 10 Evaluators, Anthropic: Claude Sonnet 4.5 is currently the preferred choice for complex coding tasks where precision is paramount. Its higher ranking suggests it is better suited for:
- Writing complex boilerplate code in new frameworks.
- Refactoring legacy systems where instruction adherence is critical.
- Debugging logic errors that require multi-step reasoning.
Anthropic: Claude Sonnet 4.6, while trailing in this specific evaluation, remains a robust option for general-purpose tasks and iterative drafting where rapid prototyping is prioritized.
Verdict
Based on our comparative evaluation, Anthropic: Claude Sonnet 4.5 outperforms the 4.6 version in the specific context of coding tasks. The 1.36 score gap suggests that for developers requiring the highest fidelity in code generation and adherence to technical instructions, the 4.5 model currently provides a more reliable output.