Argon FieldnotesAn independent Gemini guide
Menu
Release notes   Gemini 4 Argon announced. Initial access is limited.Availability snapshot ·
Developer reference

Gemini 4 Argon pricing, explained

The announced introductory and later rates, plus a simple way to budget a hypothetical workload.

Plan your token budget

Your usage. An estimated cost.

Announced rates · USD

Enter separate uncached and cached counts. Caching assumes the announced 95% discount applies and the tokens are eligible. Excludes tools, storage, infrastructure, taxes and retries. This is not a billing quote.

Token subtotal$0.400$0.800USD · token-only estimateRead the cache assumptions ↗
Token categoryIntroductoryAfter introduction
Input$2 / 1M$4 / 1M
Output$10 / 1M$20 / 1M

A hypothetical cost example

Suppose a workload uses 100,000 uncached input tokens and 20,000 output tokens. Using the introductory rates, the token subtotal is:

(100,000 ÷ 1,000,000 × $2) + (20,000 ÷ 1,000,000 × $10) = $0.40

At the later listed rates, the same token counts produce a $0.80 subtotal. These are arithmetic examples, not measured Argon sessions. They exclude tool charges, infrastructure, taxes, repeated attempts and any other fees. Confirm billing documentation before committing spend.

Four workload scenarios

These token counts are hypothetical budgets, not estimates of what Argon actually consumes. All amounts are USD token-only subtotals. The extraction batch assumes 2,000 input and 500 output tokens per request.

WorkloadUncached inputCached inputOutputIntroductoryLater
One coding attempt100,000020,000$0.400$0.800
Document synthesis50,00005,000$0.150$0.300
100 extraction requests200,000050,000$0.900$1.800
Repeated context, if cache-eligible10,00090,00020,000$0.229$0.458

Choose a use case, adapt its prompt template, and replace these counts with recorded usage when you can test it.

What the cache discount means

A 95% discount leaves 5% of the ordinary input rate: $0.10 per million cached input tokens at the introductory rate, and $0.20 at the later rate, if the announced discount applies in both periods. These are derived rates, not tested billing results.

A long repeated instruction is not automatically a cache hit. Check channel-specific eligibility, cache lifetime, minimum size and any storage charges before relying on savings. Those implementation details are not verified here. The cached scenario assumes exactly 90,000 eligible cached tokens and 10,000 uncached tokens; the other scenarios assume none.

Budget the accepted result

A task may need several requests, larger tool responses, or a retry after an unsuccessful attempt. Count those costs together. For example, if three attempts cost $0.40 each and only one result is accepted, the token cost per accepted result is $1.20. Human review, tools, infrastructure and any other charges remain separate.

Use the coding evaluation protocol to record attempted and accepted tasks. A shorter prompt can reduce input while omitting information necessary for success; judge the full workflow rather than optimizing token counts alone.

Pricing does not establish availability

Read the access snapshot before interpreting a published rate as a working API offer. This page records the announcement, not a verified account entitlement or an independently tested billing result.

Published base token prices across the comparison set

Per million tokens, USD. Sources checked October 1, 2026. These are published rates with different release and pricing-stage conditions, not a test of equal task cost.

Model / pricing stageInputOutputSource
Gemini 4 Argon (introductory)$2$10Provider
Gemini 4 Argon (later)$4$20Provider
GPT-6 Astra (starting rates)$10$50Provider
Claude Fable 5.1 (base rates)$10$50Provider

For the same uncached input and output counts, Argon's introductory token subtotal is one fifth of the listed Astra or Fable base subtotal. At Argon's later rates it is two fifths. Those ratios do not assume the models need the same tokens, retries or review work in practice.

Cache pricing needs its own assumptions

Fable's documentation lists cache reads at $0.25 per million tokens and separate cache-write rates. Argon's cached-input rate is derived from its announced 95% discount. Read Fable's pricing documentation and the Argon assumptions before comparing a cached workload. No cached billing test was performed here.

The Vals listing adds benchmark-specific cost and latency context. Keep those measurements separate from base-price arithmetic and account access.