Artificial Analysis published its evaluation of Inception's Mercury 2.5 on September 23, adding an independent set of results for a model launched earlier this month. The figures give developers a more specific picture of its tradeoffs: very fast text generation, a separate wait for the answer to begin, and a quality score below several models Inception named at launch. Artificial Analysis's changelog dates the evaluation; Inception's announcement dates the release to September 8.
The benchmark site's token rates differ from Inception's current API documentation, which advertises a launch discount. Its reported spending per evaluation task therefore needs to be read alongside the provider's current prices.
The big change
- What changed: Developers now have independent evidence for judging Mercury 2.5 alongside other text models. Inception's model emits blocks of refined text; the new evaluation measures its output speed and assesses answer quality separately.
- Why it matters: An application can receive text quickly once generation starts and still wait for an answer to begin. Buyers choosing models for interactive software need both measurements, alongside evidence that the model completes their tasks correctly.
- What to watch: Mercury exposes several reasoning settings, while the evaluation page does not identify its exact effort setting. A deployment decision therefore turns on testing the intended setting and checking the current billable rate. Both choices affect how well the public results describe the proposed use.
Output speed and answer delay measure different things
When checked on September 24, Artificial Analysis's Inception provider benchmark showed Mercury 2.5 generating 770.4 output tokens per second, with 2.95 seconds to the first answer token. Its calculated response time for 500 answer tokens was 3.60 seconds.
At that output rate, generating 500 tokens takes about 0.65 seconds. Adding the initial 2.95-second wait gives about 3.60 seconds. This reproduces the benchmark's calculation for that response length; an application's complete interaction has its own timing.
The published performance methodology uses about 10,000 input tokens and at least 1,500 answer tokens for its default workload. It reports the median over the previous 72 hours and standardizes speed measurements with a common tokenizer. Output speed measures text arriving after generation begins. Time to first answer includes the wait through any preceding reasoning.
Inception reported 1,107 tokens per second at launch, from a separate measurement whose workload and reasoning settings have not been matched to this API benchmark. Inception's streaming documentation explains that Mercury emits blocks of refined text, with an optional mode that displays intermediate refinement steps.
Quality and cost on the current index
Mercury 2.5 scores 12 on the Artificial Analysis Intelligence Index. Version 4.3.2 combines ten evaluations, weighted across agent tasks, coding, scientific reasoning and general capabilities. The methodology describes a primarily English, text-based suite. A score of 12 is an index result, not a percentage of ordinary work the model can complete.
Inception's launch comparison named GPT-5.6 Luna at low effort, Gemini 3.5 Flash-Lite and Claude Haiku 4.5. Their current AA results provide relevant context. GPT-6 Luna at low effort is included because AA identifies it as the newer replacement for the GPT-5.6 entry.
Published configuration | Current AA index | Reported cost per index task |
|---|---|---|
Mercury 2.5, reasoning; effort unspecified | 12 | $0.06 |
21 | $0.01 | |
21 | $0.0045 | |
22 | $0.12 | |
17 | $0.21 |
These are the site's rounded figures checked September 24, using its current index and cost method. Mercury's exact reasoning effort and Haiku's reasoning budget were not specified in the retrieved model records. The table compares the published configurations; it cannot establish performance with equal reasoning budgets or on a buyer's own workload.
Mercury has the lowest index result in this group. Its reported evaluation cost falls below Gemini and Haiku, but above the two Luna entries.
Promotional prices differ from the benchmark's rates
Artificial Analysis calculates cost per task from evaluation token use and token prices, including reasoning and cache costs, then weights the results by each evaluation's share of the index. Its $0.06 figure for Mercury is an average across benchmark tasks, including unsuccessful attempts.
The model page lists $0.25 per million input tokens and $0.75 per million output tokens. Inception's current model documentation instead lists standard rates of $0.20 for input and $0.75 for output, discounted to $0.04 and $0.15 respectively. Cached input is discounted from $0.02 to $0.004 per million tokens. The documentation does not give an end date for the promotion.
The published benchmark spending therefore cannot serve as a quote for the discounted API. Applying an 80% reduction to the table would also skip over the different input rate and the benchmark's cache accounting. The official documentation supplies the current unit prices; a deployment's actual token use determines its bill.
Reasoning settings remain part of the deployment decision
Inception's reasoning documentation exposes four settings for mercury-2.5: instant, low, medium and high. Medium is the API default. Artificial Analysis labels its tested model as reasoning but does not identify which of those settings produced the published Mercury entry.
A team assessing Mercury for production needs to measure answer correctness and timing with the setting it intends to use. The public evaluation does not provide a separate quality and latency result for each of the four settings.



