PeerLM logoPeerLM
All Comparisons

MiniMax: MiniMax M2.5 vs Google: Gemini 3.1 Pro Preview: Coding Performance with 10 Evaluators

In our latest Coding Performance with 10 Evaluators benchmark, we compare the efficiency of MiniMax: MiniMax M2.5 against the advanced capabilities of Google: Gemini 3.1 Pro Preview.

MiniMax: MiniMax M2.5

3.2

preference score

vs

Google: Gemini 3.1 Pro Preview

6.8

preference score

Judges ranked the responses in this Run against each other; the rank is mapped onto a 0–10 scale. It shows which response was preferred, not how good either one is — and it is not a percentage, a pass rate, or a check that the output was correct.

Sample size for this comparison was not recorded. Treat it as directional.

Evidence clarification: this article predates recorded sample provenance. Treat its conclusions as claims about the displayed examples; they do not establish general model superiority, verified correctness, or production suitability.

Key Findings

Top PerformerGoogle: Gemini 3.1 Pro Preview

Ranked #1 in overall coding performance and instruction following.

Cost AdvantageMiniMax: MiniMax M2.5

Significantly more affordable, with a total cost of $0.002185 per response.

Depth of OutputGoogle: Gemini 3.1 Pro Preview

Produced more comprehensive responses with an average of 1,612 completion tokens.

Specifications

SpecMiniMax: MiniMax M2.5Google: Gemini 3.1 Pro Preview
Providerminimaxgoogle
Context Length205K1.0M
Input Price (per 1M tokens)$0.27$2.00
Output Price (per 1M tokens)$1.08$12.00
Max Output Tokens128,00065,536
Tierstandardpremium

Our Verdict

Google: Gemini 3.1 Pro Preview is the clear winner for coding tasks, demonstrating superior accuracy and instruction following in our comparative analysis. While MiniMax: MiniMax M2.5 is remarkably cost-effective, it currently trails behind Gemini in terms of overall performance and utility for complex programming challenges.

Overview

As the landscape of Large Language Models (LLMs) evolves, developers are increasingly looking for tools that offer the best balance of coding accuracy and instruction following. In this report, we evaluate the performance of two prominent models: MiniMax: MiniMax M2.5 and Google: Gemini 3.1 Pro Preview. This assessment utilizes PeerLM's comparative evaluation framework, relying on feedback from 10 expert evaluators to determine how these models handle complex coding tasks.

Benchmark Results

Our comparative analysis shows a distinct difference in performance rankings. When evaluating Coding Performance with 10 Evaluators, the models demonstrated the following overall scores based on expert human preference:

ModelOverall ScoreRank
Google: Gemini 3.1 Pro Preview6.841
MiniMax: MiniMax M2.53.162

Criteria Breakdown

The benchmarking suite focused on two primary pillars: Accuracy and Instruction Following. Because this was a comparative evaluation, these scores reflect the relative preference of the evaluators rather than static rubrics. Gemini 3.1 Pro Preview emerged as the clear leader in both categories, indicating a higher degree of reliability when handling complex logic and intricate programming requirements.

  • Accuracy: Gemini 3.1 Pro Preview demonstrated superior precision in code generation, resulting in fewer logical errors compared to M2.5.
  • Instruction Following: Evaluators noted that Gemini was more adept at adhering to specific constraints and formatting requests within the coding prompts.

Cost & Latency

Cost efficiency remains a critical factor for enterprise developers. While Gemini 3.1 Pro Preview offers higher performance, it comes at a higher price point per request. Conversely, MiniMax M2.5 offers a highly economical alternative for projects where budget is the primary driver.

ModelTotal Cost (USD)Avg Completion Tokens
Google: Gemini 3.1 Pro Preview$0.0791061,612
MiniMax: MiniMax M2.5$0.002185427

It is important to note that the significantly higher completion token count for Google's model suggests it provides more verbose, detailed coding explanations compared to the more concise output of the MiniMax model.

Use Cases

When to choose Google: Gemini 3.1 Pro Preview

This model is best suited for complex software engineering tasks, architectural planning, and debugging large codebases where the cost of a hallucination or logic error far outweighs the cost per token. Its superior ranking in our benchmark makes it the preferred choice for mission-critical code generation.

When to choose MiniMax: MiniMax M2.5

MiniMax M2.5 is an excellent candidate for high-volume, cost-sensitive applications. If your use case involves simpler scripting, boilerplate generation, or rapid prototyping where you need to manage a high throughput of requests at a fraction of the cost, M2.5 provides a compelling value proposition.

Verdict

When comparing MiniMax: MiniMax M2.5 vs Google: Gemini 3.1 Pro Preview, the data clearly supports Gemini for high-stakes coding performance. While MiniMax offers significant cost savings, the performance delta in our 10-evaluator study highlights that Google's latest preview model is currently the more robust tool for developers demanding high accuracy and strict adherence to coding instructions.

Backed by real data

View the Full Evaluation Report

See every response, score, and evaluator judgment behind this comparison. All data from PeerLM's blind evaluation pipeline.

View Report

Run your own Monitor

Compare MiniMax: MiniMax M2.5 and Google: Gemini 3.1 Pro Preview on sampled production prompts, with frozen criteria and inspectable evidence.

Start a Monitor

Get a free managed report

We'll run a full evaluation with your real prompts and deliver a detailed recommendation. Free for qualified teams.

Request Report

Methodology

Evaluated using PeerLM's blind evaluation pipeline with 4 responses per model across 2 criteria.