The published snapshot
Selected percentages from Google DeepMind’s model page, rechecked October 1, 2026. The source is Google’s comparison table, not an independent Argon Fieldnotes experiment.
| Benchmark | Gemini 4 Argon | Claude Opus 5.5 |
|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.2% |
| FrontierSWE v2 | 55.0% | 62.3% |
| Terminal-bench 4.0 | 57.4% | 66.4% |
| CWE-bench v1 | 68.0% | 67.0% |
DeepSWE: Argon +3.7 percentage points. FrontierSWE: Opus +7.3. Terminal-bench: Opus +9.0. CWE-bench: Argon +1.0. These are absolute percentage-point gaps.
Interpret the result sources
Google’s methodology document describes using different result sources for individual benchmarks. Read those settings before treating the pair as a same-environment test. The displayed values do not include the human corrections, latency or total spending for your task.
A coding result also does not establish research or drafting quality. If those are your use cases, evaluate citation support and source fidelity directly with the same documents rather than transferring a coding ranking to a different task.
Choose the test that matches your work
For tool-heavy development, evaluate the exact shell, repository and permissions your workflow requires. For repository changes, verify the proposed patch and inspect unrelated edits. For document work, use the source-bounded research prompt and check every cited conclusion.
Track accepted results and review time for each candidate. Our use-case guides provide task criteria and pricing scenarios provide provisional Argon token budgets. They are preparation tools, not claims that either model has passed a test here.