Include both a narrow task and one that requires reading across files. Avoid evaluating only examples whose answer is already in the prompt. Write down what would count as a correct result before the run begins.
Keep the environment comparable
Record the exact model version, instruction text, repository revision, available tools, time limit and spending limit. Use a clean working copy for each model or attempt. If one run gets extra feedback, tool access or retries, record that difference rather than describing the results as equivalent.
Give the agent only the permissions required for the task. Keep credentials and production access outside the evaluation environment. Review proposed changes before they are merged or deployed.
Verify beyond the agent’s summary
- Run the original failing check against the proposed change.
- Run the nearby checks that exercise the affected behavior.
- Inspect the diff for unrelated edits or a weakened test.
- Check whether the explanation matches the actual code change.
- Record remaining review work and every unsuccessful attempt.
Score accepted results
Use a simple result sheet: pass or fail, review minutes, elapsed time, total cost and any permission violations. Separate correct-but-expensive outcomes from incorrect ones. A model that produces a passing result after substantial human repair should not receive the same result as one whose change passes without repair.
This is an evaluation method you can use when access permits. Argon Fieldnotes has not performed this test on Argon.
The published benchmark snapshot can help you choose an evaluation area. Your own repository tests answer the more specific adoption question.
Turn the protocol into an instruction
Use our coding and debugging prompt to specify the failure, permitted files, allowed checks and stopping conditions. Supply the real test command and repository revision rather than assuming the model knows them.
A good first pilot is a reproducible bug with a small expected diff. Defer broad migrations until you have enough successful bounded tasks to understand review effort and failure modes. Compare candidates using the same task and published evidence, then track costs using the pricing scenarios.