# Gemini 4 Argon launches to cyber testers; API access still has no date
> Google announced Gemini 4 Argon for trusted cyber testers, with paid API and AI Ultra access planned next but undated. Vals AI's first independent scores lead on some tasks and trail on others; the public API catalog lists no Argon endpoint.
By BIG CHANGE Editorial
Published: 2026-09-30T23:56:07.213Z
Updated: 2026-10-01T02:16:11.971Z
Canonical: https://bigchange.ai/blog/gemini-4-argon-restricted-access-independent-benchmarks

AI-generated editorial illustration by BIG CHANGE.
Google announced **Gemini 4 Argon** on September 30 with strong scores for knowledge work and coding, a claimed million-token output limit, and prices for a future API release. Google is rolling the model out to an initial group of trusted cybersecurity defenders and testers. It has not given developers a public model ID or a date when paid API customers or Google AI Ultra subscribers can use it. [Google's announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/) and the [Gemini API model catalog](https://ai.google.dev/gemini-api/docs/models) show the gap between the launch and access for most readers.
## The big change
- **What changed:** Google has put a new frontier model into selected cyber defenders' hands while most developers and subscribers still cannot access it. Its headline capability and price claims are now public ahead of a public API.
- **Why it matters:** Independent Vals AI evaluations place Argon first on broad knowledge work and a finance agent test, so teams comparing models have a serious new candidate. The uneven results across tasks and restricted rollout leave their own workflow results and costs unresolved.
- **What to watch:** Google has named paid API customers and AI Ultra subscribers as the next groups, without a date. A documented model ID and service limits will let developers test the work that matters to them, including whether the claimed million-token output can be used in practice.
## Who can use Gemini 4 now
Google says the first recipients are trusted cyber defenders and testers in its [Fairwind program](https://blog.google/innovation-and-ai/technology/safety-security/fairwind-program/). Fairwind is a limited-access program for selected Google Cloud customers, government agencies and security partners. Google says it is giving trusted defenders Argon without its usual cyber guardrails so they can use its defensive capabilities. The initial release is a controlled test of a particularly sensitive use case.
Google plans to expand to developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers. It has published no date for that step. As checked on October 1, Google's [Gemini API model directory](https://ai.google.dev/gemini-api/docs/models) listed Gemini 3 models and no Gemini 4 Argon identifier. Its [API pricing page](https://ai.google.dev/gemini-api/docs/pricing) also had no Argon row. An API call recipe would therefore require inventing a model name and access path that Google has not documented.
The company nevertheless stated future token rates in its launch post. The introductory price is **$2 per million input tokens and $10 per million output tokens**, with cached input priced 95% below the input rate. After the introductory period, Google says the rates will rise to **$4 and $20**. It did not state when the introductory period ends. These are announced rates, not an available self-service API price today. The rate per token also says little about the cost of a long agent task until the amount of reasoning, tool use and output for that task is known.
## Strong scores, with visible weak spots
[Vals AI's independent Argon record](https://www.vals.ai/models/google_gemini-4-argon) reports **68.90% ± 0.97** on the Vals Index, first among 41 models in its table. It places Argon first of 73 on Finance Agent v2 at **65.40% ± 0.32**. On Vibe Code Bench it ranks second of 106 at **91.91% ± 1.90**; on Code Migration, second of 71 at **68.17% ± 4.35**; and on CyberBench v1.1, second of 43 at **77.86% ± 5.30**. Those results support a strong early showing across several tasks, without making Argon the best model on each one.
The same evaluator ranks Argon fifth of 42 on Terminal-Bench 4.0 at **57.58% ± 2.31**, fifteenth of 106 on MedScribe, and seventh of eight on its computer-use CUA-bench at **4.83%**. Its measured cost per test also varies sharply: **$15.68** for a Vals Index test and **$193.78** for CUA-bench. These are evaluator task costs, separate from Google's announced prices per million tokens. Vals says some agent evaluations use a fixed harness while others use a model's native agent; its reported standard errors do not capture variation across prompts, seeds or deployment settings. A small score difference should be read in that context.
[Google DeepMind's own comparison table](https://deepmind.google/models/gemini/) adds more counterexamples to any claim of an across-the-board lead. It shows Argon ahead of GPT-6 Astra on DeepSWE v1.1, **77.9% to 74.1%**, but behind Astra on FrontierSWE v2, **55.0% to 65.5%**, and Terminal-Bench Science, **57.6% to 68.1%**. Claude Opus 5.5 leads Argon on Terminal-bench 4.0, **66.4% to 57.4%**, and PostTrainBench, **49.3% to 45.3%**. On the table's offline OSWorld-2.0 subset, Astra scores **72.6%** against Argon's **69.2%**. Those comparisons use Google's selected table and its stated benchmark setups; they are not a substitute for customer testing in a documented public service.
Google says Argon can generate up to **one million output tokens in a single trajectory**, up from a previous 64,000-token limit. Vals lists a one-million-token *context* window but configured its Argon evaluation with a **262,144-token maximum output**. That setting does not establish a hard limit on Google's future product. It does mean the public Vals results do not demonstrate a full million-token generation, and Google has not yet published a public API entry that resolves the offered output limit for ordinary developers.
## Internal uses and safety claims need their own evidence
Google describes Argon agents helping with C/C++ to Rust migrations, a quantum-computing subroutine and data-center memory optimization. It says one set of memory changes freed more than 300 TiB after rollout. It also says a partner, Wiz, found a critical vulnerability affecting healthcare software with Argon. These are Google's accounts of internal or partner work; the launch materials do not provide enough independent detail to verify those outcomes or reproduce them. Google says the large Rust rewrites still undergo automated and manual review before production.
For safety, Google describes refusal training for harmful requests, defenses against indirect prompt injection, monitoring of model reasoning and actions for behavior outside user intent, and stronger isolation of evaluation environments. Its assertion that Argon is its most resilient model against indirect prompt injection is a vendor claim. The initial cyber cohort is also receiving access without standard cyber guardrails, according to Google, which makes the scope of access and monitoring especially consequential as it widens.
There is conflicting early evidence about day-to-day usefulness. [Axios reported](https://www.axios.com/2026/09/30/google-gemini-4) that Bloomberg, citing anonymous sources, said some Google employees found Argon's internal performance lacking; Axios also reported that Google told Bloomberg the account was inaccurate. Google told Axios that employees had used the model for difficult coding and research work. Neither account settles how well a public version will perform. The useful next evidence is access under documented settings, accompanied by task results and costs that others can examine.
## Sources
- [Google's Gemini 4 Argon announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/) — Official launch, access plan, future prices, claimed output limit, internal projects and safety statements. Internal outcomes and safety comparisons are vendor claims.
- [Google DeepMind Gemini model page](https://deepmind.google/models/gemini/) — Google's selected head-to-head benchmark table, including Argon leads and losses; checked October 1.
- [Gemini API model catalog](https://ai.google.dev/gemini-api/docs/models) — Public developer model catalog; no Gemini 4 Argon model ID listed when checked. This does not rule out private access.
- [Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing) — Public developer pricing page; no Argon row when checked. Launch post contains announced future rates.
- [Vals AI Gemini 4 Argon evaluation](https://www.vals.ai/models/google_gemini-4-argon) — Independent evaluator's benchmark scores, ranks, task costs and evaluation configuration, including 262,144 maximum output tokens.
- [Vals AI evaluation methodology](https://www.vals.ai/methodology) — Explains task design, agent harness variation and what its standard errors do and do not capture.
- [Google Fairwind program](https://blog.google/innovation-and-ai/technology/safety-security/fairwind-program/) — Google's background on the limited cyber defense program and intended partners.
- [Axios launch reporting](https://www.axios.com/2026/09/30/google-gemini-4) — Independent access context and Google comment; relays Bloomberg's anonymous employee skepticism and Google's denial.
- [Ars Technica launch reporting](https://arstechnica.com/google/2026/09/google-announces-gemini-4-argon-ai-model-but-you-cant-use-it-yet/) — Independent launch context; internal uses remain Google-reported claims.
The BIG CHANGE newsletter
The big picture. At your pace.
Recent stories on AI and robotics, the shifts worth watching and practical ideas to use. Choose a daily briefing, weekly digest or monthly perspective.
Sent at 09:00 Belgrade time: daily, Mondays or the first of the month. Your first edition arrives at the next scheduled send after you confirm.
Your privacy, your choice.
Necessary storage supports site security and remembers your choices. Optional Google Analytics stays off until you allow it. You can read every story with necessary storage only. Privacy details