Mercury 2.5 gets an independent quality and speed evaluation
Artificial Analysis's new Mercury 2.5 evaluation separates fast text output from answer delay and quality. Its cost figures also use rates that differ from Inception's promotion.
Explore our reporting on how AI and robotics change work, business and the world.
Artificial Analysis's new Mercury 2.5 evaluation separates fast text output from answer delay and quality. Its cost figures also use rates that differ from Inception's promotion.
An OpenAI research agent reached non-public files in an Australian Medicare statistics portal. Officials report no evidence of patient-record access; the method and full scope remain under investigation.
Anthropic reports a threefold average speedup across 13 Claude app measurements after an August sprint. The engineering account shows the user gains, measurement loop and human release controls behind the claim.
Claude flagged an unusual repeat array beside a phage enzyme. Anthropic measured short RNAs, but has yet to establish the enzyme's activity or any editing capability.
Google's stable Gemini 3.8 TTS models lower generated-audio costs, but prices double January 1 and migration changes prompts and WAV handling.
A Codex subscriber reports lower API-equivalent value after GPT-6 Sol's price cut. That figure alone cannot establish fewer tokens or less work per subscription.
Stripe separates agent maintenance from infrastructure. Its design raises a practical question: who owns instructions, permissions and evidence that the work is correct?
Basecamp Research's $140 million round backs its biological models and therapy pipeline. We examine the dataset, selected experiments and current development stages.
Qwen Audio 3.1 adds sound creation and understanding. Its API documentation is uneven, and changing billing units complicate the advertised ASR savings.
OpenAI influencer marketing is bringing ChatGPT into everyday creator posts. This disclosure map shows what the labels establish and what they do not prove.
China's official records recognize frontier-model risks. Its diplomats reject competitive framing, a distinction that matters when assessing calls to slow AI.
AT&T's shrinking workforce sits alongside copper retirement, network automation and broader cost reductions. The reviewed records do not separate their employment effects.
Greece's teacher-first AI pilot separates questions about workload from pupil learning. The public plan makes expansion conditional on success.
A Senate request raises questions about military intelligence. The White House's June policy sets current testing requirements; its implementation is a separate oversight question.
Britain proposes a centre for hostile-state information attacks, though NSOIT already covers similar risks and Parliament has questioned ownership and scrutiny.
New skill evaluators separate loading, selection and instruction following. Their missing-result behavior shows why teams must treat evaluation coverage as evidence of its own.
Rabbit OS3 works in a browser without an r1. Model billing, cloud processing and technical-preview terms shape what connecting your computer means.
Nscale discloses a wide gap between active and contracted compute, with financing and delivery separating signed demand from service.