Mercury 2.5 gets an independent quality and speed evaluation
Artificial Analysis's new Mercury 2.5 evaluation separates fast text output from answer delay and quality. Its cost figures also use rates that differ from Inception's promotion.
Explore our reporting on how AI and robotics change work, business and the world.
Artificial Analysis's new Mercury 2.5 evaluation separates fast text output from answer delay and quality. Its cost figures also use rates that differ from Inception's promotion.
Google's stable Gemini 3.8 TTS models lower generated-audio costs, but prices double January 1 and migration changes prompts and WAV handling.
A Codex subscriber reports lower API-equivalent value after GPT-6 Sol's price cut. That figure alone cannot establish fewer tokens or less work per subscription.
Basecamp Research's $140 million round backs its biological models and therapy pipeline. We examine the dataset, selected experiments and current development stages.
Qwen Audio 3.1 adds sound creation and understanding. Its API documentation is uneven, and changing billing units complicate the advertised ASR savings.
Rabbit OS3 works in a browser without an r1. Model billing, cloud processing and technical-preview terms shape what connecting your computer means.
OpenAI adds a cache dashboard, diagnostics and explicit breakpoints to help developers trace repeated input costs and control which context gets cached.
Eight models, four tiers: compare intelligence scores, API prices and measured task costs for the latest OpenAI and Claude releases.
Grok 4.7’s launch promises stronger coding and office work. We examine Matthew Berman’s coverage, independent tests and what teams should measure before switching.
TypeSafe CEO Diogo Almeida argues for AI built around small decisions inside software. We examine Jev’s launch, its limits and what would prove a meaningful change.