Overview
In this technical breakdown, we evaluate the coding capabilities of two leading lightweight models: Anthropic: Claude Haiku 4.5 and Google: Gemini 3 Flash Preview. By utilizing PeerLM's rigorous comparative evaluation framework, we assessed these models across a suite focused on Coding Performance with 10 Evaluators. The result provides clear insight into which model delivers superior coding logic and instruction adherence in real-world scenarios.
Benchmark Results
The comparative evaluation focused on ranking-based performance, where 10 specialized evaluators assessed the models on their ability to generate accurate, functional code while following complex instructions. Google: Gemini 3 Flash Preview emerged as the top-performing model in this specific benchmark run.
| Model | Rank | Overall Score | Accuracy | Instruction Following |
|---|---|---|---|---|
| Google: Gemini 3 Flash Preview | 1 | 7.95 | 7.95 | 7.95 |
| Anthropic: Claude Haiku 4.5 | 2 | 2.05 | 2.05 | 2.05 |
Criteria Breakdown
The evaluation centered on two primary pillars: Accuracy and Instruction Following. In the context of Coding Performance with 10 Evaluators, accuracy measures the syntax correctness and logical soundness of the generated code snippets. Instruction Following measures how well the model adheres to specific constraints, such as library requirements, style guides, or API usage patterns.
Google: Gemini 3 Flash Preview demonstrated a significant lead with a score of 7.95, indicating higher reliability in complex coding tasks. Anthropic: Claude Haiku 4.5 scored 2.05, placing it behind the current leader in this specific test suite.
Cost & Latency
Efficiency is a critical factor for developers integrating LLMs into code generation pipelines. Below is the cost breakdown for the evaluation run:
- Google: Gemini 3 Flash Preview: Total cost of $0.002085, with a cost per output token of $0.003791.
- Anthropic: Claude Haiku 4.5: Total cost of $0.004878, with a cost per output token of $0.006206.
Beyond the raw metrics, Google: Gemini 3 Flash Preview proved to be the more cost-effective solution within this specific evaluation subset, processing requests with lower resource expenditure while maintaining higher performance scores.
Use Cases
Given the results of the Coding Performance with 10 Evaluators benchmark, these models are suited for different implementation strategies:
- Google: Gemini 3 Flash Preview: Best for high-volume automated code generation, complex refactoring tasks, and assistant-based coding tools where accuracy and instruction adherence are paramount.
- Anthropic: Claude Haiku 4.5: While trailing in this specific coding benchmark, it remains a viable candidate for lightweight, latency-sensitive tasks where the specific coding requirements of this evaluation suite may not be the primary driver.
Verdict
For developers prioritizing coding performance, the Anthropic: Claude Haiku 4.5 vs Google: Gemini 3 Flash Preview comparison clearly favors Google. With an overall score of 7.95 versus 2.05, Gemini 3 Flash Preview is the clear winner for coding-centric workflows. Furthermore, its superior cost efficiency makes it a compelling choice for production-grade AI engineering.