Google announced Gemini 4 Argon on September 30, 2026, at 20:00 UTC, targeting software engineering, enterprise knowledge work and cyber defense. Initial access goes to selected cyber defenders through Fairwind; wider access is planned, beginning with paid API customers and Google AI Ultra subscribers.
For application builders, the useful work now is to prepare representative evaluations and a budget that survives the introductory offer. Keep the existing production model until access and workload performance are established.
Who can access Argon?
Cyber-defense teams: Apply through the official Fairwind program if your organization fits its remit. Google reviews applicants and conducts organizational checks; an application does not guarantee admission or immediate access.
General API builders: The announcement gives no general-access date. Prepare evaluation inputs, but do not commit a migration date or assume a usable public model identifier.
Consumer subscribers: Google AI Ultra is part of the planned wider rollout, not confirmation that every Ultra account can use Argon today. Wait for account-specific availability.
Fairwind’s access rules restrict use to internal cybersecurity, incident-response or penetration-testing teams and prohibit sharing, redistribution or resale. Participating organizations must track employee use and apply authentication and access controls. It is not a route for powering arbitrary customer-facing applications.
The Gemini API model catalog and pricing page checked on September 30 did not list Argon. That is a documentation snapshot, not proof that private endpoints do not exist. The rates below come from the announcement.
Budget for the price after the introductory offer
Google’s announcement and pricing footnote specify the following USD rates per million tokens. No introductory expiry date is given.
Token category | Introductory | After introductory period |
|---|---|---|
Uncached input | $2 | $4 |
Output | $10 | $20 |
Cached input, calculated | $0.10 | $0.20 only if the discount continues |
The cached-input row is arithmetic: Google announces a 95% discount from the input price. It is not a separately verified API rate-card entry, and the later cached rate assumes that discount persists. Cache eligibility, creation and storage charges still need confirmation.
Illustrative calculation, not a model test: Assume aggregate usage across requests of 100,000 input tokens and 50,000 billable output tokens. Hold that usage fixed and exclude all other charges. This is not a claim about what fits in one request.
No cached input: The introductory token subtotal is (0.1 × $2) + (0.05 × $10) = $0.70.
90% of input qualifies for cache reads: The subtotal becomes (0.01 × $2) + (0.09 × $0.10) + (0.05 × $10) = $0.529.
That is about 24.4% less, not a 95% reduction in the job’s cost: output still costs $0.50 in both cases. At the later base rates, the no-cache subtotal is $1.40; with the same cache mix and an unchanged discount, it is $1.058.
Use the later prices for a durable budget and treat introductory savings as temporary. Add tools, retries, infrastructure and human review separately, then divide total spending by accepted tasks. The Opus 5.5 budgeting guide explores the same distinction between cheaper repeated context and cheaper completed work; its billing rules should not be transferred to Argon.
A million output tokens is a ceiling, not a task budget
Google specifies an output limit of one million tokens. That is not an input-context specification. At the announced output rates, one million billable output tokens alone would cost $10 initially or $20 later, before input and other costs.
The announcement does not settle how reasoning tokens relate to visible output and that ceiling, or how long such a generation would take. Plan explicit time and spending limits for each task. For long jobs, save intermediate work outside the model and define what counts as completion; validate interruption and recovery behavior once the actual interface is documented.
Use benchmark results to design the evaluation
The CWE-bench v1 leaderboard reports Argon with Antigravity at 68% programmatic pass@1, tied for first, versus 62% under its separate judge panel. The benchmark uses 120 held-out tasks, four rollouts per task at high reasoning and a one-hour limit per rollout.
Those are benchmark-provider results, not RohitAI measurements. The graders apply different checks, and the leaderboard compares model-and-harness combinations rather than model weights alone. Neither percentage is a forecast for your repositories.
For a patching evaluation, require both a repaired security property and preserved existing functionality. A plausible explanation or a patch that passes only one narrow check is insufficient. Match the time allowance to your workflow; a one-hour benchmark allowance does not establish interactive latency.
What to prepare before wider access
Freeze representative tasks. Save starting repositories or source documents, expected outcomes and failure cases. Include incomplete evidence, incorrect patches and interrupted runs. Establish results for your current production baseline.
Define the acceptance decision. Set correctness, elapsed-time and spending limits before running Argon. Record rejected attempts and review effort as well as successful outputs; otherwise retries can hide the true cost.
Confirm the serving details. Before integration, obtain the documented model identifier, account eligibility, regions, quotas, supported tools and output formats, input limits, billing categories and data terms. Do not infer these from another Gemini model.
Run a bounded comparison once authorized. Keep tools, permissions and task inputs controlled. Record the served model, route, settings, tokens, elapsed time and review outcome. Expand only if it meets the acceptance criteria within the later-price budget.
Until those conditions are met, prepare the comparison and keep the existing production route. The announcement supplies enough information to plan; it does not yet supply the evidence for a general migration.
Methodology: AI-assisted reporting and analysis based on Google’s announcement, Fairwind policy, public developer documentation and benchmark-provider reporting, checked September 30, 2026. Cost examples are arithmetic under stated assumptions. No hands-on Argon testing was performed.
