PeerLM logoPeerLM
All Comparisons

OpenAI: o3 vs Google: Gemini 3.1 Pro Preview: Coding Performance with 10 Evaluators

In our latest Coding Performance with 10 Evaluators benchmark, we compare OpenAI: o3 and Google: Gemini 3.1 Pro Preview to see which model leads in software development tasks.

OpenAI: o3

4.7

preference score

vs

Google: Gemini 3.1 Pro Preview

5.3

preference score

Judges ranked the responses in this Run against each other; the rank is mapped onto a 0–10 scale. It shows which response was preferred, not how good either one is — and it is not a percentage, a pass rate, or a check that the output was correct.

Sample size for this comparison was not recorded. Treat it as directional.

Evidence clarification: this article predates recorded sample provenance. Treat its conclusions as claims about the displayed examples; they do not establish general model superiority, verified correctness, or production suitability.

Key Findings

Top PerformanceGoogle: Gemini 3.1 Pro Preview

Secured the highest overall score of 5.26 in coding accuracy and instruction following.

Best ValueOpenAI: o3

Delivered high-quality results at a significantly lower total cost of $0.026432.

ThroughputGoogle: Gemini 3.1 Pro Preview

Generated more comprehensive code blocks with an average completion length of 1,612 tokens.

Specifications

SpecOpenAI: o3Google: Gemini 3.1 Pro Preview
Provideropenaigoogle
Context Length200K1.0M
Input Price (per 1M tokens)$2.00$2.00
Output Price (per 1M tokens)$8.00$12.00
Max Output Tokens100,00065,536
Tierpremiumpremium

Our Verdict

Google: Gemini 3.1 Pro Preview outperformed OpenAI: o3 in terms of raw accuracy and instruction following, making it the preferred choice for complex coding tasks. However, OpenAI: o3 offers superior cost-efficiency, proving that it is a highly viable alternative for teams optimizing for budget without sacrificing significant performance.

Overview

As the demand for AI-driven software engineering tools grows, choosing the right model is critical. In this evaluation, we compare OpenAI: o3 vs Google: Gemini 3.1 Pro Preview specifically regarding their coding performance. Using a panel of 10 expert evaluators, we assessed how these models handle complex coding prompts, instruction adherence, and overall output accuracy.

Benchmark Results

The comparative evaluation highlights a clear leader in terms of raw performance, while also showcasing distinct advantages for each model in different deployment scenarios.

ModelOverall ScoreAccuracyInstruction Following
Google: Gemini 3.1 Pro Preview5.265.265.26
OpenAI: o34.744.744.74

Criteria Breakdown

The evaluation focused on two primary pillars: Accuracy and Instruction Following. In coding contexts, accuracy refers to the syntactical correctness and logical soundness of the generated code, while instruction following measures how well the model adheres to specific architectural constraints or developer preferences.

Google: Gemini 3.1 Pro Preview demonstrated superior performance across both metrics, securing an overall score of 5.26. OpenAI: o3 followed closely with a score of 4.74. While the score spread of 0.52 indicates a competitive landscape, the evaluators consistently preferred the depth of output provided by the Gemini variant.

Cost & Latency

For engineering teams, the trade-off between performance and cost is paramount. Below is a breakdown of the economic impact of using these models for coding tasks.

  • Google: Gemini 3.1 Pro Preview: Total cost of $0.079106 across evaluated responses, with an average completion length of 1,612 tokens.
  • OpenAI: o3: Total cost of $0.026432 across evaluated responses, with an average completion length of 772 tokens.

OpenAI: o3 stands out as the more cost-effective solution, providing a high level of performance at roughly one-third the cost per task compared to Gemini 3.1 Pro Preview. This makes it an ideal candidate for high-volume coding environments where budget efficiency is a priority.

Use Cases

Google: Gemini 3.1 Pro Preview

Best suited for complex, multi-file code generation and architectural design tasks where the model's ability to maintain context over longer outputs (averaging 1,612 completion tokens) results in higher-quality, ready-to-run solutions.

OpenAI: o3

Ideally positioned for iterative development, unit test generation, and smaller snippet implementations where rapid, reliable, and cost-efficient coding support is required.

Verdict

Google: Gemini 3.1 Pro Preview is the current performance leader for coding tasks, providing higher accuracy and better instruction adherence in our 10-evaluator study. However, OpenAI: o3 remains a formidable competitor, offering significant cost savings that make it a highly attractive option for scalable development workflows.

Backed by real data

View the Full Evaluation Report

See every response, score, and evaluator judgment behind this comparison. All data from PeerLM's blind evaluation pipeline.

View Report

Run your own Monitor

Compare OpenAI: o3 and Google: Gemini 3.1 Pro Preview on sampled production prompts, with frozen criteria and inspectable evidence.

Start a Monitor

Get a free managed report

We'll run a full evaluation with your real prompts and deliver a detailed recommendation. Free for qualified teams.

Request Report

Methodology

Evaluated using PeerLM's blind evaluation pipeline with 4 responses per model across 2 criteria.