AI Tools

Gemini 4 Argon: Release Status, API Pricing, Benchmarks and the 1M Output Limit

Gemini 4 Argon explained: who can use it now, $2/$10 intro pricing vs $4/$20 later, Google's benchmark results, and what the 1M output limit costs.

Photo de profil de Long Nguyen

Long Nguyen

Développeur fullstack · Ingénieur IA · Chercheur

• • 6 min de lecture •

What is Gemini 4 Argon?

Gemini 4 Argon is the first model of Google's Gemini 4 generation, announced by Google DeepMind on . It sits above the Gemini 3.8 line as Google's new frontier model and is aimed at long-running, multi-step work rather than quick chat: real-world software engineering, enterprise knowledge work such as legal and finance, and defensive cybersecurity.

Three things set it apart from the previous generation, according to Google's official Gemini 4 Argon announcement:

  • A 1M-token output limit, up from 64K on the previous models. This is output, not context: the model can generate hundreds of thousands of tokens in a single response.
  • Frontier cyber-defense capability: it can find, validate and patch software vulnerabilities on its own, which is why the rollout is gated.
  • Introductory API pricing of $2 / $10 per million input / output tokens, doubling to $4 / $20 after the introductory period.

The short version for anyone planning work around it: the model is announced, priced and benchmarked, but you cannot call it from the public API yet.

Is Gemini 4 available yet?

Not for most people. Google is releasing Argon in phases and has not published a date for general availability. As of , access looks like this:

Stage Who gets access Status
1. Fairwind Program A cohort of trusted cyber defenders, who receive the model without cyber guardrails Rolling out now
2. Government pre-release review The U.S. government's voluntary pre-release model access process In progress
3. First paid wave Paid Gemini API customers and Google AI Ultra subscribers "As soon as possible", no date
4. Broad release Developers, enterprises and consumers generally After guardrail iteration, no date

Two practical consequences. First, the announcement does not name an API model ID, so any code sample you see online with a specific Gemini 4 model string is a guess. Second, if you are on a free tier or a non-Ultra consumer plan, expect to wait behind paying API customers.

Gemini 4 Argon API pricing

Google published the price before the model is generally available, which is useful for budgeting now:

Token type Introductory price (per 1M) After introductory period (per 1M)
Input $2.00 $4.00
Cached input $0.10 (95% off input) Not stated; $0.20 if the 95% discount carries over
Output $10.00 $20.00

Google has not said how long the introductory period lasts. Plan your margins on the $4 / $20 price, and treat the lower rate as a temporary discount rather than the real cost of the model.

What a typical agent call costs

Output tokens dominate the bill for long-horizon work, and that matters more on a model that can now write up to 1M of them. A worked example for one agentic coding step with a 200K-token prompt, of which 150K is a cached repository context, producing 40K tokens of output:

Line item Introductory Post-introductory
50K uncached input $0.10 $0.20
150K cached input $0.015 $0.03
40K output $0.40 $0.80
Total per call $0.515 $1.03

Output is about 78% of that total. A single call that uses the full 1M output limit costs $10 at the introductory price and $20 afterwards, before any input. Prompt caching cuts the input side hard, but nothing discounts output, so the lever that actually controls cost is how much you let the model write.

Gemini 4 Argon benchmarks

These are the scores and rankings Google reports in its launch post. Where Google gave a ranking but no number, the table says so rather than filling one in.

Benchmark What it measures Argon result (Google-reported)
DeepSWE v1.1 Real-world, long-horizon software engineering tasks 77.9%, new state of the art
Vals Index Economic impact across finance, coding, legal and tax work, weighted by U.S. GDP contribution #1 (score not given in the post)
AutomationBench (Zapier) End-to-end execution of core business functions 51.3%, ranked #1
LVBench Long video understanding 91.7%, state of the art
CWE-bench v1 Remediating security vulnerabilities 68%, tied for first
Vals Finance Agent v2 Multi-step financial research Leading (score not given)
Harvey Legal Agent Benchmark Legal research and drafting Leading (score not given)
Gray Swan IPI Robustness to indirect prompt injection Leading (score not given)

How much weight to put on these numbers

All of the above are vendor-reported, on a model almost nobody outside Google can test yet. That does not make them wrong, but it means there is no independent replication. Bloomberg reported internal skepticism at Google about how well the model performs in areas such as coding, despite the benchmark lead. The sensible reading: the long-output and agentic-workflow numbers point to a real capability jump, and your own evaluation on your own tasks is still the only result that should decide a migration.

What the 1M output token limit actually changes

Most developers have been working around output caps for years: split a large refactor into dozens of calls, generate a long report section by section, stitch everything back together, then fix the seams. A 1M output ceiling, roughly 15 times the previous 64K, removes much of that orchestration for single-pass jobs.

Google's own internal examples show what that looks like at scale:

  • Codebase migration: Argon agents are moving C/C++ code to Rust inside Google, from core libraries such as re2 up to 800K+ lines for the Fuchsia Zircon kernel, with automated and manual audits before anything reaches production.
  • Performance work: on libgav1, Google's open-source video decoder, agents replaced 32K lines of SIMD code in an existing Rust port. The result is memory-safe and runs 2.7x faster than that port, with identical output.
  • Infrastructure: agents analyzing fleet-wide profiling data applied memory optimizations that freed more than 300 TiB across Google's data centers.

The trade-offs nobody mentions in the launch coverage

  • A failed long generation is expensive. If a 600K-token response dies at token 500K because of a timeout or a dropped connection, you have paid for 500K output tokens and have nothing usable. Stream responses, and design jobs so partial output can be checkpointed or resumed.
  • Latency scales with output. A response that is hundreds of thousands of tokens long takes minutes, not seconds. That suits batch jobs and background agents, not anything a user waits on in a UI.
  • Review effort scales too. One pass that produces 800K lines of migrated code is only useful if you can verify it. Google itself runs emulation testing and manual review on these rewrites. If your pipeline has no tests to catch regressions, a bigger output limit just produces bigger unverified diffs.
  • Always set an explicit output cap. The ceiling is a maximum, not a target. Size max_output_tokens to the job so a model that rambles cannot turn a $0.50 call into a $10 one.

Why Gemini 4 access is restricted: cyber capability and safeguards

The gated rollout is a direct result of what the model can do. Google trained Argon to autonomously find, validate and patch critical vulnerabilities, and it is giving trusted defenders and its own teams a version without cyber guardrails. Wiz is already using it in its Scan for Good program for public infrastructure, where Google says the model found a critical flaw exposing personal data in healthcare software used by hospitals, which earlier frontier models had missed.

Before broad release, Google says it is hardening four areas:

Area What Google is doing Why it matters if you build on it
Misuse Refusing cyber and CBRN attack requests, with monitoring of the model's internal activations to spot misuse Legitimate security tooling may hit refusals on the public version that Fairwind users do not see
Prompt injection Adversarial training and automated red teaming; Google claims the lead on Gray Swan's IPI benchmark Lower, not zero, risk for agents that read web pages, emails or uploaded files
Misalignment Monitors watch chain-of-thought and actions and can stop execution when the model goes beyond user intent Long agent runs may be halted mid-task; handle that as a normal outcome
System hardening Sandboxes isolated and sealed before high-risk training or evaluation A reminder to sandbox your own agents that execute code or touch production

How to prepare your app for Gemini 4 now

You cannot call Argon yet, but you can make the switch a configuration change instead of a rewrite when it arrives.

  1. Move the model name into configuration. Never hardcode it. When Google publishes the Argon model ID, you change one environment variable.
  2. Build a small evaluation set from your real tasks. Twenty to fifty representative inputs with known-good outputs is enough to compare your current model against Argon on day one, instead of trusting launch benchmarks.
  3. Restructure prompts for caching. Put the stable part first (system instructions, schemas, repository or document context) and the variable part last. At 95% off cached input, this is the cheapest optimization available.
  4. Cap output and set billing alerts. Decide the largest output each job genuinely needs, and set a budget alert before the first Argon call, not after the first invoice.
  5. Price your product on post-introductory rates. If a feature only works at $2 / $10, it does not work.

A model-agnostic call with the official Google Gen AI Python SDK looks like this. The model name comes from the environment, and output is capped per job:

import os
from google import genai
from google.genai import types

client = genai.Client()  # reads GEMINI_API_KEY from the environment

MODEL = os.environ["GEMINI_MODEL"]  # set to the Argon ID once Google publishes it

def run_job(prompt: str, max_out: int = 32_000) -> str:
    response = client.models.generate_content(
        model=MODEL,
        contents=prompt,
        config=types.GenerateContentConfig(max_output_tokens=max_out),
    )
    return response.text

If your team is weighing whether Argon justifies rebuilding a chatbot, a report-generation agent or an internal automation, it helps to separate what the model changes from what your pipeline needs anyway. That is the kind of decision our AI agent and chatbot development work starts with: evaluation on your own data first, model choice second.

Should you switch to Gemini 4 Argon?

Once it is available, it depends on the shape of your workload more than on the benchmark table:

Your workload Recommendation Reason
Large refactors, code migrations, long agentic coding runs Test early The 1M output limit and DeepSWE result target exactly this
Long reports, contract or financial analysis from many documents Test early Single-pass long output removes section-stitching; legal and finance results are its strongest claims
Long video or chart-heavy document understanding Test LVBench 91.7% is the headline multimodal result
High-volume, short responses (support chat, classification, extraction) Stay on a Flash-class model You pay frontier prices for capability you do not use; latency matters more
Security testing and vulnerability work Apply via Fairwind if eligible The unguardrailed cyber capability is only available there
Anything with a launch date in the next few weeks Do not depend on it No general-availability date or public model ID yet

The pattern most teams will end up with is routing, not replacement: a fast, cheap model for the bulk of requests and Argon for the long, hard jobs where a single high-quality pass saves more than it costs.

If you want a second opinion on where Argon fits in your stack, or help setting up the evaluation and routing before access opens, request a free quote and describe the workload. There is no cost or account needed to ask.

FAQ

Questions fréquentes

When was Gemini 4 released?

Google announced Gemini 4 Argon on September 30, 2026. It is a phased release: it is rolling out first to trusted cyber defenders through the Fairwind Program, and Google has not given a date for general availability.

Can I use Gemini 4 Argon in the Gemini API today?

Not yet for most users. Google says broader access will start with paid Gemini API customers and Google AI Ultra subscribers, but no date or public model ID had been published as of October 1, 2026.

How much does Gemini 4 Argon cost?

The introductory API price is $2 per million input tokens and $10 per million output tokens, with cached input at 95% off the input price. After the introductory period, the price rises to $4 per million input and $20 per million output. Google has not said how long the introductory period lasts.

What is the Gemini 4 output token limit?

Gemini 4 Argon can generate up to 1 million output tokens in a single response, up from 64K on the previous generation. That is an output limit, separate from how much input context the model accepts.

What is Gemini 4 Argon best at?

Google positions it for long-horizon software engineering, enterprise knowledge work such as legal and financial research, long video understanding, and defensive cybersecurity. It reports 77.9% on DeepSWE v1.1 and the top position on the Vals Index, although these results are self-reported and not yet independently replicated.

What is the Fairwind Program?

Fairwind is Google DeepMind's program for trusted cyber defenders. Participants are the first to get Gemini 4 Argon, in a version without cyber guardrails, so they can use its vulnerability discovery and patching capabilities.

Restez informé avec Netalith

Recevez des ressources de développement, des mises à jour produit et des offres spéciales directement dans votre boîte mail.