# Mercury 2.5 gets an independent quality and speed evaluation

> Artificial Analysis's new Mercury 2.5 evaluation separates fast text output from answer delay and quality. Its cost figures also use rates that differ from Inception's promotion.

By BIG CHANGE Editorial

Published: 2026-09-24T02:45:45.727Z
Updated: 2026-09-24T02:45:45.727Z
Canonical: https://bigchange.ai/blog/mercury-2-5-independent-evaluation

![Conceptual charcoal illustration of an analog stopwatch on a ledge beside receding, unbranded computing cabinets.](https://bigchange.ai/api/media/file/mercury-evaluation-stopwatch-hero-v1.png)
AI-generated conceptual illustration by BIG CHANGE.

Artificial Analysis published its evaluation of Inception's Mercury 2.5 on September 23, adding an independent set of results for a model launched earlier this month. The figures give developers a more specific picture of its tradeoffs: very fast text generation, a separate wait for the answer to begin, and a quality score below several models Inception named at launch. [Artificial Analysis's changelog](https://artificialanalysis.ai/changelog) dates the evaluation; [Inception's announcement](https://www.inceptionlabs.ai/blog/introducing-mercury-2-5) dates the release to September 8.

The benchmark site's token rates differ from Inception's current API documentation, which advertises a launch discount. Its reported spending per evaluation task therefore needs to be read alongside the provider's current prices.

## The big change

- **What changed:** Developers now have independent evidence for judging Mercury 2.5 alongside other text models. Inception's model emits blocks of refined text; the new evaluation measures its output speed and assesses answer quality separately.
- **Why it matters:** An application can receive text quickly once generation starts and still wait for an answer to begin. Buyers choosing models for interactive software need both measurements, alongside evidence that the model completes their tasks correctly.
- **What to watch:** Mercury exposes several reasoning settings, while the evaluation page does not identify its exact effort setting. A deployment decision therefore turns on testing the intended setting and checking the current billable rate. Both choices affect how well the public results describe the proposed use.

## Output speed and answer delay measure different things

When checked on September 24, Artificial Analysis's [Inception provider benchmark](https://artificialanalysis.ai/models/mercury-2-5/providers) showed Mercury 2.5 generating 770.4 output tokens per second, with 2.95 seconds to the first answer token. Its calculated response time for 500 answer tokens was 3.60 seconds.

At that output rate, generating 500 tokens takes about 0.65 seconds. Adding the initial 2.95-second wait gives about 3.60 seconds. This reproduces the benchmark's calculation for that response length; an application's complete interaction has its own timing.

The [published performance methodology](https://artificialanalysis.ai/methodology/performance-benchmarking) uses about 10,000 input tokens and at least 1,500 answer tokens for its default workload. It reports the median over the previous 72 hours and standardizes speed measurements with a common tokenizer. Output speed measures text arriving after generation begins. Time to first answer includes the wait through any preceding reasoning.

Inception reported 1,107 tokens per second at launch, from a separate measurement whose workload and reasoning settings have not been matched to this API benchmark. Inception's [streaming documentation](https://docs.inceptionlabs.ai/capabilities/streaming) explains that Mercury emits blocks of refined text, with an optional mode that displays intermediate refinement steps.

## Quality and cost on the current index

Mercury 2.5 scores 12 on the [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/models/mercury-2-5). Version 4.3.2 combines ten evaluations, weighted across agent tasks, coding, scientific reasoning and general capabilities. The [methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) describes a primarily English, text-based suite. A score of 12 is an index result, not a percentage of ordinary work the model can complete.

Inception's launch comparison named GPT-5.6 Luna at low effort, Gemini 3.5 Flash-Lite and Claude Haiku 4.5. Their current AA results provide relevant context. GPT-6 Luna at low effort is included because AA identifies it as the newer replacement for the GPT-5.6 entry.

| Published configuration | Current AA index | Reported cost per index task |
| --- | --- | --- |
| Mercury 2.5, reasoning; effort unspecified | 12 | $0.06 |
| [GPT-5.6 Luna, low effort](https://artificialanalysis.ai/models/gpt-5-6-luna-low) | 21 | $0.01 |
| [GPT-6 Luna, low effort](https://artificialanalysis.ai/models/gpt-6-luna-low) | 21 | $0.0045 |
| [Gemini 3.5 Flash-Lite, high effort](https://artificialanalysis.ai/models/gemini-3-5-flash-lite) | 22 | $0.12 |
| [Claude Haiku 4.5, reasoning](https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning) | 17 | $0.21 |

These are the site's rounded figures checked September 24, using its current index and cost method. Mercury's exact reasoning effort and Haiku's reasoning budget were not specified in the retrieved model records. The table compares the published configurations; it cannot establish performance with equal reasoning budgets or on a buyer's own workload.

Mercury has the lowest index result in this group. Its reported evaluation cost falls below Gemini and Haiku, but above the two Luna entries.

## Promotional prices differ from the benchmark's rates

Artificial Analysis calculates cost per task from evaluation token use and token prices, including reasoning and cache costs, then weights the results by each evaluation's share of the index. Its $0.06 figure for Mercury is an average across benchmark tasks, including unsuccessful attempts.

The model page lists $0.25 per million input tokens and $0.75 per million output tokens. Inception's [current model documentation](https://docs.inceptionlabs.ai/get-started/models) instead lists standard rates of $0.20 for input and $0.75 for output, discounted to $0.04 and $0.15 respectively. Cached input is discounted from $0.02 to $0.004 per million tokens. The documentation does not give an end date for the promotion.

The published benchmark spending therefore cannot serve as a quote for the discounted API. Applying an 80% reduction to the table would also skip over the different input rate and the benchmark's cache accounting. The official documentation supplies the current unit prices; a deployment's actual token use determines its bill.

## Reasoning settings remain part of the deployment decision

Inception's [reasoning documentation](https://docs.inceptionlabs.ai/capabilities/reasoning-efforts) exposes four settings for `mercury-2.5`: instant, low, medium and high. Medium is the API default. Artificial Analysis labels its tested model as reasoning but does not identify which of those settings produced the published Mercury entry.

A team assessing Mercury for production needs to measure answer correctness and timing with the setting it intends to use. The public evaluation does not provide a separate quality and latency result for each of the four settings.

## Sources

- [Artificial Analysis changelog](https://artificialanalysis.ai/changelog) — The September 23, 2026 entry records the new Mercury 2.5 intelligence evaluation and Inception endpoint performance. This establishes the news date, distinct from the September 8 model release.
- [Artificial Analysis: Mercury 2.5 model evaluation](https://artificialanalysis.ai/models/mercury-2-5) — Checked September 24, 2026. Reports index 12 and $0.06 weighted cost per index task. Labels the model as reasoning without specifying effort. Its token rates differ from Inception's current documentation.
- [Artificial Analysis: Mercury 2.5 provider benchmark](https://artificialanalysis.ai/models/mercury-2-5/providers) — Direct page retrieval on September 24, 2026 reports Inception at 770.4 output tokens/s, 2.95 seconds to first answer and 3.60 calculated seconds for 500 answer tokens. These are changing benchmark measurements, not application guarantees.
- [Artificial Analysis: API performance methodology](https://artificialanalysis.ai/methodology/performance-benchmarking) — Checked September 24, 2026. Defines the default 10,000-input-token workload, at least 1,500 answer tokens, rolling 72-hour median, common speed tokenizer and separate first-answer and output-speed metrics. General parameters do not reveal Mercury's exact reasoning effort.
- [Artificial Analysis: Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) — Checked September 24, 2026. Version 4.3.2 combines ten evaluations across four weighted categories. The suite is primarily English and text based; its composite score does not describe every workload.
- [Inception: Introducing Mercury 2.5](https://www.inceptionlabs.ai/blog/introducing-mercury-2-5) — Original September 8 launch, including Inception's 1,107 tokens/s claim and named comparison models. Vendor claims and customer examples are not results of the September 23 independent evaluation.
- [Inception: Models and pricing](https://docs.inceptionlabs.ai/get-started/models) — Checked September 24, 2026. Lists Mercury 2.5 at $0.20 input and $0.75 output per million tokens, discounted to $0.04 and $0.15; cached input is $0.004 during the promotion. No promotion end date was specified.
- [Inception: Reasoning efforts](https://docs.inceptionlabs.ai/capabilities/reasoning-efforts) — Checked September 24, 2026. Defines instant, low, medium and high for mercury-2.5, with medium as the API default. This does not establish which setting Artificial Analysis used.
- [Inception: Streaming and diffusion](https://docs.inceptionlabs.ai/capabilities/streaming) — Checked September 24, 2026. Documents output in refined blocks and the optional display of intermediate denoising steps. It explains API behavior, not independently measured speed.
- [Artificial Analysis: GPT-5.6 Luna at low effort](https://artificialanalysis.ai/models/gpt-5-6-luna-low) — Checked September 24, 2026. Current index 21 and reported task cost $0.01. Artificial Analysis marks this entry as superseded by GPT-6 Luna low; that label does not establish the provider's retirement policy.
- [Artificial Analysis: GPT-6 Luna at low effort](https://artificialanalysis.ai/models/gpt-6-luna-low) — Checked September 24, 2026. Current index 21 and reported task cost $0.0045 for the low-effort configuration, using the same current index and cost method as the comparison table.
- [Artificial Analysis: Gemini 3.5 Flash-Lite](https://artificialanalysis.ai/models/gemini-3-5-flash-lite) — Checked September 24, 2026. Current index 22 and reported task cost $0.12. The page's embedded model record identifies high reasoning effort; the comparison retains that setting.
- [Artificial Analysis: Claude Haiku 4.5 with reasoning](https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning) — Checked September 24, 2026. Current index 17 and reported task cost $0.21 for the reasoning variant. The retrieved record does not specify its reasoning-token budget.
