PeerLM logoPeerLM
All Comparisons

Anthropic: Claude Fable 5.1 vs Qwen: Qwen3.8 Max (0902): Coding Performance with 10 Evaluators

A comparative analysis of Anthropic: Claude Fable 5.1 and Qwen: Qwen3.8 Max (0902) based on Coding Performance with 10 Evaluators.

Anthropic: Claude Fable 5.1

5.3

preference score

vs

Qwen: Qwen3.8 Max (0902)

4.7

preference score

Judges ranked the responses in this Run against each other; the rank is mapped onto a 0–10 scale. It shows which response was preferred, not how good either one is — and it is not a percentage, a pass rate, or a check that the output was correct.

4
test cases
at least 4
evaluated responses per model
9
judge models (of 10 seated)
36
usable judgments
Run
Sep 16, 2026

Evaluated responses: Qwen: Qwen3.8 Max (0902): 4 · Anthropic: Claude Fable 5.1: 4

Task scope: Coding Performance with 10 Evaluators

Small sample: 4 responses per model. This shows which output judges preferred on these examples — not which model is better in general, and not a check that the code or facts were correct.

Key Findings

Preference ScoreAnthropic: Claude Fable 5.1

Achieved a higher relative preference score of 5.28.

LatencyAnthropic: Claude Fable 5.1

Faster response times with an average latency of 2879ms.

Cost EfficiencyQwen: Qwen3.8 Max (0902)

Significantly more cost-effective for these coding tasks.

Specifications

SpecAnthropic: Claude Fable 5.1Qwen: Qwen3.8 Max (0902)
Provideranthropicqwen
Context Length1.0M1.0M
Input Price (per 1M tokens)$10.00$2.00
Output Price (per 1M tokens)$50.00$6.00
Max Output Tokens128,000131,072
Tierfrontieradvanced

Our Verdict

Anthropic: Claude Fable 5.1 performed better in terms of relative judge preference and speed on these specific coding examples. However, Qwen: Qwen3.8 Max (0902) offers a notable advantage in cost efficiency. Because the sample size is limited to 4 test cases, these results serve as a directional comparison rather than a definitive ranking.

Overview

This report provides a comparative analysis of two prominent large language models, Anthropic: Claude Fable 5.1 and Qwen: Qwen3.8 Max (0902), focusing specifically on their Coding Performance with 10 Evaluators. The evaluation was conducted using PeerLM's comparative framework, which utilizes relative preference ranking rather than static rubric scoring. By analyzing model responses side-by-side, we captured the nuanced preferences of 9 independent judge models across 4 distinct coding test cases.

Benchmark Results

In this evaluation, each model generated 4 responses, resulting in a total of 36 individual judgments provided by the judge panel. The scoring represents the relative preference rank mapped onto a 0-10 scale, indicating how often a model's output was favored by the judges during the evaluation process.

ModelOverall ScoreAvg Latency (ms)Total Cost (USD)
Anthropic: Claude Fable 5.15.2828790.15854
Qwen: Qwen3.8 Max (0902)4.7237540.021068

Criteria Breakdown

The evaluation centered on a holistic assessment of coding output. It is important to note that the scores presented here reflect the aggregate relative preference of the judge models. Because this was a comparative ranking exercise, the scores represent the models' ability to satisfy the judges' expectations relative to one another within the scope of Coding Performance with 10 Evaluators.

Cost & Latency

Performance in a coding environment often necessitates a balance between speed and cost. Anthropic: Claude Fable 5.1 demonstrated a faster average latency of 2879ms compared to the 3754ms recorded for Qwen: Qwen3.8 Max (0902). However, Qwen: Qwen3.8 Max (0902) proved significantly more cost-efficient for these tasks, with a total cost of 0.021068 USD compared to 0.15854 USD for Claude Fable 5.1. Developers should weigh these latency and budget considerations based on their specific integration requirements.

Use Cases

The models were tested on their ability to generate code snippets and follow specific programming instructions. These results are most applicable to developers seeking to understand how these models perform in a comparative setting. Given the nature of the evaluation, these models are best suited for tasks where coding assistance is required, though users should perform their own validation before deploying generated code in sensitive environments.

Limitations

This evaluation is based on a limited sample size of 4 test cases per model, resulting in 4 responses per model. These findings are directional and represent a snapshot of relative performance in a specific task scope. This evaluation did not verify code execution, check for logical correctness against ground truth, or assess performance in production-grade software engineering pipelines.

Verdict

Based on the relative preference rankings, Anthropic: Claude Fable 5.1 outperformed Qwen: Qwen3.8 Max (0902) on these specific examples, achieving a higher aggregate score. While Claude Fable 5.1 offers lower latency, Qwen: Qwen3.8 Max (0902) provides a more economical option for similar coding tasks. Users should consider these results as a comparative baseline rather than an absolute measure of model quality.

Backed by real data

View the Full Evaluation Report

See every response, score, and evaluator judgment behind this comparison. All data from PeerLM's blind evaluation pipeline.

View Report

Run your own Monitor

Compare Anthropic: Claude Fable 5.1 and Qwen: Qwen3.8 Max (0902) on sampled production prompts, with frozen criteria and inspectable evidence.

Start a Monitor

Get a free managed report

We'll run a full evaluation with your real prompts and deliver a detailed recommendation. Free for qualified teams.

Request Report

Methodology

Blind ranking evaluation: judges ranked 4 responses per model against each other. Scores express relative preference on these responses, not an absolute quality rating.