Coding and security results in Google's snapshot
All four rows below use Google's model table, checked October 1, 2026. They are reported values, not our measurements.
| Benchmark | Gemini 4 Argon | Claude Fable 5.1 |
|---|---|---|
| DeepSWE v1.1 | 77.9% | 67.4% |
| FrontierSWE v2 | 55.0% | 56.3% |
| Terminal-bench 4.0 | 57.4% | 57.9% |
| CWE-bench v1 | 68.0% | 58.0% |
Argon is ahead on displayed DeepSWE and CWE-bench values; Fable is ahead on FrontierSWE and Terminal-bench. Keep those tasks separate. Anthropic publishes its own Fable Terminal-Bench result with different evaluation conditions; we retain Google's values consistently rather than combining sources into one ranking.
Read Google's methodology and Anthropic's evaluation notes before interpreting gaps. A published comparison is a starting point for a test in your own environment.
Context and output are different limits
Google announces a 1M output-token ceiling for Argon. Anthropic's Fable documentation lists 1M context and 128K maximum output. These describe different dimensions; a large input window does not imply an equally large output allowance. Argon announcement · Fable specifications.
For Argon, the long-output guide also preserves the difference between the announced ceiling and Vals' evaluation setting. We have not run either model at its maximum output limit.
Compare the published rates, then the whole task
Argon's announced introductory rates are $2 input / $10 output per million tokens. Fable's documented base rates are $10 / $50, with cache reads at $0.25 per million tokens. Cache creation, eligibility and other charges matter; a cache-read rate alone is not a total task price. Argon rates · Fable rates.
Use the dated price reference for arithmetic, then record repeated attempts and review time. Do not infer equal token usage or identical cache behavior across models.
A practical selection sequence
- Confirm which model and channel your organization can actually use.
- Choose a bounded task and hold inputs, tool permissions and budgets constant.
- Run an independent acceptance check and log refusals, retries and corrections.
- Compare cost per accepted result and completion time before expanding the workload.
Fable's published availability can make it an actionable candidate sooner. Argon's displayed task results can justify testing it when access permits. These are editorial selection criteria, not a universal recommendation or a claim that Fable is safer overall.