| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% |
| Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% |
| CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% |
Selected rows, not the full evaluation suite. Same source page for all displayed values. Results are not independently verified by this publication.
For a focused decision, see Argon vs GPT-6 Astra or Argon vs Claude Opus 5.5.
What the methodology changes
Google’s evaluation methodology PDF says results generally use a single attempt and high thinking settings. Some Argon results are computed by Google, while comparison values may come from provider reports or public leaderboards. DeepSWE uses a mini-swe harness for Argon.
That means this table is a record of published evidence rather than a controlled experiment run by one independent evaluator. Before making a fine-grained comparison, inspect the individual benchmark’s task definitions, harness, resource budget and reporting rules.
Read each row as its own question
Argon has the highest displayed DeepSWE value here, while Astra has the highest FrontierSWE value and Opus the highest Terminal-bench value. CWE-bench shows a tie between Argon and Astra at the displayed precision. The pattern does not support declaring one model the universal winner.
Use a relevant benchmark to identify a candidate, then use a workflow-specific evaluation to check whether it solves your actual problem. An aggregate result cannot show the quality of a particular patch, the human review required, or your account’s eventual billing.
What we would require for an independent comparison
We would publish the task set, starting revisions, prompts, model versions, tools, budgets, attempt counts and acceptance criteria, together with the raw outcomes. Until those runs exist, this site will continue to label this material as provider-published evidence.
Business automation and long-video understanding
These additional task results come from Google's model table and keep the reported model versions intact.
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|
| AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% |
| LVBench | 91.7% | 87.5% | 79.7% | 83.7% |
AutomationBench concerns workflow execution; LVBench concerns long-video understanding. These rows do not establish citation quality, permissions compliance or factual consistency across a large generated document. Use a relevant test to shortlist candidates, then evaluate your own task.
A separate third-party view: Vals Index
Vals' Argon listing, checked October 1, 2026, displays 68.90% accuracy with ±0.97, $15.68 cost per Vals Index test, and 46 min 33 s latency. These belong to that benchmark listing; they are not our measurements, an API response-time guarantee or a quote for your workload.
The displayed ± value is preserved without assigning it a statistical interpretation that the listing does not explain here. The hyperparameter panel lists high reasoning effort and 262,144 maximum output tokens, and notes that some benchmarks use different settings. That helps explain why a third-party run should remain separate from the provider snapshot.
Google announces a 1M output ceiling; Vals lists a smaller evaluation setting. See the source distinction. Do not use the benchmark latency to claim that every Argon request takes 46 minutes.