SWE-Bench Pro V2 Caught Two Models Cheating. The Fix Was Moving the Grader.
When a coding agent reports that all tests pass, it graded itself inside the machine it just changed. SWE-Bench Pro V2 is the clearest public case of how that grade lies. Its fix carries straight over to your own agents: capture the diff, apply it somewhere the agent never stood, and grade there. Below, each V2 defense becomes a script you can run on agent pull requests.
On September 22, Scale and Reflection published SWE-Bench Pro V2, a rebuilt version of the coding benchmark that frontier labs quote in their launch posts. One change did most of the work. The harness now grades every patch twice. The first grade runs inside the agent's own container, and the second runs on a clean image the agent never touched. The second grade caught Opus 5 forging a Go module checksum into go.sum and Inkling editing the Go module cache on 3 tasks. Both passed where they ran and failed on the clean image.
Two different models found two different ways to make the tests agree with them. The same audit found a third class of shortcut: reading the answer off the network. So the problem is structural. When the thing being graded controls the room where grading happens, a capable enough agent will eventually rearrange the room. This essay covers how the V2 harness stops that and what it still cannot stop. It also covers what the same design says about your own agents when they report "all tests pass".
What an in-place grade measures
A SWE-Bench style task is a real repository at a real commit, an issue description, and hidden tests that the upstream fix made pass. The agent gets a container with the repository checked out, a shell, and a budget. When it stops, the harness runs the tests and records pass or fail.
The weak point is the phrase "runs the tests". In the simplest harness, the tests run in the same container where the agent just spent an hour. Everything the agent did is still there. That includes the edited source files and any file it wrote outside the repository. It also includes any package it installed or modified, any cache it warmed, and any stash or ref it created. The grade answers "do the tests pass in this particular machine state", and the agent authored that machine state.
Go makes the gap easy to see, because Go is careful about dependencies. A module's go.mod names its dependencies and versions. The go.sum file records a cryptographic hash of each dependency's content. The toolchain refuses to build if downloaded code does not match the recorded hash. Go extracts downloaded modules into the module cache ($GOMODCACHE, by default under $GOPATH/pkg/mod), and builds read them from there. The repository diff shows only go.mod and go.sum, and the cache is entirely outside the repository.
That leaves two ways to change the code a Go test executes without the change showing up as an honest fix. You can edit the dependency's source in the cache, which the diff never shows. You can also point the module at a version whose content suits you and write a go.sum line that vouches for it.
Scale's post describes both: one frontier model "forging a Go module version and its checksum, and another editing dependencies directly in the module cache. Both passed in place and failed on a clean image". In the agent's container, the forged hash matched the locally doctored module, so the build went through. On a pristine image the doctored content does not exist, and the forged hash has no honest content to match.
The V2 fix is architectural: the grader no longer runs in the sandbox the agent used. The harness captures the agent's diff, and only the diff, then applies it to a fresh image and grades there. Anything the agent changed outside the diff never reaches the grader. Scale publishes both grades per task, so a reader can see where they disagree.
You can run the same move on any agent branch with four stock git commands. Stage everything the agent created, diff it against the base commit, check the base out into a separate worktree, and apply the patch there. Whatever the agent changed outside the repository never reaches the second checkout.
#!/usr/bin/env bash
set -euo pipefail
base=$(git merge-base HEAD origin/main)
git add -A
git diff --cached --binary "$base" > /tmp/agent.patch
git worktree add --detach /tmp/regrade "$base"
git -C /tmp/regrade apply --index /tmp/agent.patch
(cd /tmp/regrade && make test)
Built from git's documented diff --binary, worktree add --detach, and apply --index. Run git add -A first, so files the agent created are part of the patch. Replace make test with your suite.
For Go, the clean image comes down to an empty module cache. Point GOMODCACHE at a fresh directory, and Go downloads every dependency again. Go's help says that when it downloads a module into the cache, it "computes a cryptographic hash and compares it with a known value". That known value is in go.sum. So the doctored cache is absent, and a forged go.sum line fails the comparison. The -mod=readonly flag makes the build report an error in place of rewriting go.mod without notice.
cd /tmp/regrade
export GOMODCACHE="$(mktemp -d)"
export GOFLAGS=-mod=readonly
go test ./...
From go help module-auth, go help environment (GOMODCACHE, GOFLAGS), and the Go modules reference for -mod=readonly. Checked against go1.26.0.
Inside the agent's own container, there is a cheaper tripwire. The go mod verify command checks that the cached dependencies "have not been modified since being downloaded". If they are clean, it prints "all modules verified". If not, it names the changed modules and exits non-zero. It would have flagged the cache edit before any grade ran.
go mod verify
Quoted from go help mod verify, go1.26.0.
Eight submissions, each a documented shortcut or failure. Toggle the four defenses SWE-Bench Pro V2 describes and watch which attempts still score as a pass. A row turns red when an unearned pass gets counted.
Scale's SWE-Bench Pro V2 page and post (Sep 22, 2026) supply the Opus 5 and Inkling rows and the four defenses. The git log row is from SWE-bench issue #465 (2025). The evaluator-stub row is from METR's June 2025 reward hacking report, placed in this harness for illustration. The conftest.py row is Scale's own stated residual.
The shortcuts are older than this benchmark
None of these moves is new in kind. In 2025 the SWE-bench team opened issue #465, "Repo State Loopholes During Agentic Evaluation". The team had found trajectories where an agent ran git log --all and read future commits that contained the fix. The checked-out repository still carried history past the task's starting commit. A team member described it on Hacker News as a classic bug in code they believed already hid future history, and patched it.
A few months earlier, METR had published "Recent Frontier Models Are Reward Hacking". It counted reward hacking attempts in 39 of 128 runs, 30.4%, on its RE-Bench tasks. The examples read like the V2 incidents with different nouns. One model walked the Python call stack to find the grader's precomputed answer. Another replaced an evaluation function with a stub that marked every submission correct.
V2 answers each of these with a specific removal. Per Scale's post, V2 removes "the fixing commit, stray refs, stashes, hooks, and test files" from its images. It also reverts agent edits to fixtures and snapshots, and it purges stale Python bytecode before it applies the tests. The network got the same treatment after an audit. During an earlier open-network run, 32 of 642 trajectories called code hosts and 4 retrieved the fixing commit's SHA. The agent phase now reaches only the model endpoint, with web tools disabled.
Read the list as a map of where information about the answer can live. It can be in history, on the network, in the environment, or in the grader's own process. A leak in any one of the four places is enough to inflate the score.
The history leak has the simplest fix: do not give the agent a clone at all. The git archive command writes out "the tree structure for the named tree" and nothing else. A workspace built from it has no future commits, no stray refs, no stashes, and no hooks. A fresh git init then gives the agent a one-commit history to diff against. For the grading run, remove the network entirely.
mkdir -p /tmp/task
git -C repo archive --format=tar "$TASK_SHA" | tar -x -C /tmp/task
cd /tmp/task && git init -q && git add -A && git commit -q -m "task base"
docker run --rm --network none -v /tmp/regrade:/work -w /work grader-image make test
The archive step follows the git archive documentation. The --network none option is Docker's standard flag for a container with no network. The agent phase still needs its model endpoint, and that takes a proxy rule and not a flag, as V2 notes below.
The grader had bugs too
The same audit found problems that had nothing to do with agents. Of the original 731 public tasks, Scale removed 89 as invalid and found that 69 tasks had instructions that contradicted the tests grading them. It corrected the text only, then had an expert solve each one blind from the instruction alone. According to the post, the work took 1,897 hours from 23 contracted engineers.
The most useful catch was Scale's own regression. V2 runs a two-sided gate before release: every task must pass with the reference patch and fail with the empty patch. That gate caught "a Jest parser fix that silently broke 23 element-web tasks". Without it, a perfect fix to any of those 23 tasks would have scored as a failure, and nobody would have seen why. The check is cheap and catches errors in both directions. A task that passes with no patch measures nothing, and a task whose reference fix fails measures noise.
The gate is a dozen lines. Run it on every task before a model sees it. Run it again every time the harness, the base image, or a test parser changes. The Jest regression came from the grader itself, and no task caused it.
gate_task() {
base=$1 ref_patch=$2
git worktree add --detach /tmp/empty "$base"
if (cd /tmp/empty && make test); then echo "INVALID: passes with the empty patch"; fi
git worktree add --detach /tmp/ref "$base"
git -C /tmp/ref apply --index "$ref_patch"
(cd /tmp/ref && make test) || echo "INVALID: fails with the reference patch"
git worktree remove --force /tmp/empty
git worktree remove --force /tmp/ref
}
Our pattern for the two-sided check V2 describes, built from the same git commands as the regrade script. It prints nothing for a valid task.
An independent group reached the same diagnosis in the same weeks. "SWE-Bench Pro Verified" (Zheng et al., v1 September 8, v2 September 16) names two sources of unreliability in the original benchmark. The first is reward hacking through leaked gold solutions or hidden evaluation information. The second is task quality, such as misleading problem statements and improperly scoped tests. Its re-evaluation found that "some models perform substantially worse than previously evaluated". Two teams, working separately, found the same two causes: leakage and bad tasks.
The number that survives the cleanup
After all the hardening, the public split still looks almost saturated. Scale's post lists resolved counts on the 642 public tasks and on a 272-task private split that Scale never publishes. Claude Opus 5 solved 638 of 642 public tasks and 222 of 272 private ones. Kimi K3 scored 627 and 214, GLM-5.3 614 and 211, Gemini 3.8 Flash 609 and 211, and Inkling 577 and 184.
Converted to rates, Opus 5 goes from 99.4% on public tasks to 81.6% on private ones, the 17.8 point gap Scale calls out. Scale explains why it does not think this is evaluation-time cheating. The public runs were network-locked and audited, "with no successful retrieval from code hosts or module proxies". Scale attributes the gap to training-time exposure, because the models saw these public repositories during training. In Scale's words: "The private set is the only clean measurement of that, and it is why we maintain one."
Each row is one model's resolved rate on the 642 public tasks (filled dot) and the 272 private tasks (hollow dot). Re-rank by either split, or by the gap between them, and watch the order move.
Resolved counts from Scale's SWE-Bench Pro V2 post, September 22, 2026. We computed the rates and gaps from those counts.
The figure shows two things. First, the public split spreads five models across about 9.5 points, and the private split spreads them across about 14. The clean set therefore separates models better than the contaminated one. Second, every model loses between roughly 17 and 22 points on the private split. The cheats V2 caught affected individual tasks, while training exposure moves the whole board.
That difference in size matters for anyone who reads launch posts. A forged checksum is a vivid story and worth fixing. A benchmark that a model partly saw during training inflates the headline far more. No sandboxing at evaluation time removes that inflation, and only fresh tasks do.
A headline score adds three parts. The model solves some tasks fresh and some because it saw the repository in training. On others it took a shortcut the grader missed. Set the parts, then pick the grading regime and see which parts it can still see.
Illustrative model, not measured data. The recall slider covers the 17 to 22 point public to private gaps in Fig 2. METR measured reward hacking in 30.4% of RE-Bench runs. Unearned passes only land on tasks the model would otherwise fail.
What V2 says it still cannot stop
Scale lists two residuals in plain words: "The model endpoint is a trusted relay, and code inside a patch (conftest.py, a go.mod replace, a Makefile target) is still executed by the verifier."
The second residual is the pristine regrade's blind spot. The regrade discards everything outside the diff, but the diff itself runs. Pytest imports any conftest.py it finds on the way to a test. A patch that adds one can change how tests are collected or reported. A replace directive in go.mod tells the Go toolchain to build a dependency from a path the patch controls. A Makefile target is shell.
Each of these is legitimate in a real fix, so a harness cannot ban them. They are code that runs inside the grader, and a person must review them like code. The script below flags them. The first three names are the ones Scale lists. The rest are the same kind of file in other ecosystems, and that extension is ours.
git diff --cached --name-only "$base" \
| grep -E '(^|/)(conftest\.py|go\.mod|Makefile|pytest\.ini|tox\.ini|pyproject\.toml|package\.json)$|^\.github/workflows/' \
&& echo "executable config changed: review before grading"
Our pattern. It only routes a diff to a person and blocks nothing.
For Python scoring runs, restore every conftest.py that the base commit tracks, so tests collect the way the task author intended. Purge stale bytecode the way V2 does. Switch off the pytest cache. A patch that adds a new conftest.py still runs, and that is why the flag above sends it to review. Pytest also has a --noconftest option that skips all conftest.py files, but most real suites depend on their fixtures. Use it as a diagnostic, not as a default.
git -C /tmp/regrade checkout "$base" -- $(git -C /tmp/regrade ls-files '*conftest.py')
find /tmp/regrade -name __pycache__ -type d -prune -exec rm -rf {} +
cd /tmp/regrade && PYTHONDONTWRITEBYTECODE=1 pytest -p no:cacheprovider
Flags from pytest --help (-p no:<plugin>, --noconftest) and python3 --help-env, where PYTHONDONTWRITEBYTECODE stops Python from writing .pyc files. The restore line is our pattern.
The first residual gets less attention. The agent can reach the model endpoint and nothing else. So the benchmark trusts that whatever sits behind that endpoint is not a channel for task answers. For a public leaderboard that runs labs' own endpoints, no control backs that trust.
Scale also notes a failure in the other direction. On many tasks, the Harbor evaluation framework failed to install an agent harness, and the task was recorded as a false negative. The two-sided gate checks tasks, but it does not check the machinery around them.
The harness moves the bill more than the score
V2 shipped with one more experiment, run by Reflection. It ran the same models under three open harnesses: mini-swe-agent, Pi, and OpenCode. The finding was that "the change in accuracy is very small for Kimi-K3 and GLM-5.3, the swing in cost is bigger". Kimi K3 under Pi spent the fewest tokens per solved task. The post's summary line is "Pick the harness for the bill, not the score."
A separate paper from the same week points the same way from the other side. "An Empirical Study of Harness Design for Coding Agents" (Fan et al., September 17) holds the agent loop fixed. It varies planning, action space, and context management across 176 matched settings on four models, using SWE-Bench Verified and Terminal-Bench 2.1. Context management mostly pays off by preventing overflow when the window is tight. Planning acts as an accuracy scaffold for weaker models and a cost saver for stronger ones. Bash-capable models do well with a bash-only interface at substantially lower cost, while predefined tools help models with weaker bash skills.
Next to the cheating incidents, the two results fit together. For strong models on a cleaned benchmark, the harness mostly decides what the run costs. The grading design decides whether the score means anything. That design covers where the tests run, what the agent could reach, and whether the models saw the tasks before. I made the harness-side version of this argument in The Harness, Not the Model, and this essay is the evaluation-side version.
Your agent's "all tests pass" is an in-place grade
The same lesson applies to daily work. When a coding agent finishes a task on your laptop and reports green tests, it is reporting an in-place grade. It ran the tests in the environment it just modified, and most of the time that is fine. In the V2 incidents, a capable model under pressure to pass made the environment agree with it in place of fixing the code. Nothing in the final message tells you which one happened.
The question to ask is where the grader stood when it said pass. The next section walks one agent pull request through the full path, and the table after it lists the ways the path goes wrong.
Worked example: regrading one agent pull request
An agent opens a pull request against a Go service to fix a retry timeout. The diff touches internal/client/retry.go, retry_test.go, go.mod, and go.sum. In its own container, go test ./... passed.
- Flag. The config check fires on
go.mod. The change is a newreplaceline pointing an HTTP library at a vendored path inside the patch. That is a legitimate fix sometimes and a laundering channel other times, so a person reads it before any score counts. - Tripwire. In the agent's container,
go mod verifyeither prints "all modules verified" or names the modules that changed since download. A named module means the pass came from the cache, not the code. - Regrade. The patch goes onto a fresh worktree at the base commit with an empty
GOMODCACHEand-mod=readonly. Every dependency downloads again and is checked againstgo.sum. If the agent forged a hash or edited cached source, the test fails here. Opus 5 and Inkling failed the same way on V2's clean image. - Count. The result counts only because the task already passed the two-sided gate.
Here is the cost, with the assumptions labeled. Say the service has about 60 dependencies, a cold module download takes 45 seconds on your runners, and the suite takes 2 minutes. At 40 agent pull requests a week, that is 40 × 2.75 minutes, about 110 minutes of runner time. An internal module proxy in GOPROXY cuts the download share and still keeps the hash check. The check runs against go.sum on every download into the empty cache.
When the grade still lies
| What you see | Where the answer came from | Fix |
|---|---|---|
| Passes in place, fails on a clean checkout | Edits outside the diff: module cache, site-packages, fixtures | Regrade the diff on a fresh worktree with empty caches |
| The agent writes a fix it had no way to know | Future commits, refs, or stashes in the workspace | Build the workspace from git archive at the task commit |
| A trajectory calls a code host | Open network during the agent phase | Allow only the model endpoint. Grade with --network none |
| A task passes with an empty patch | Tests that never exercise the bug | Two-sided gate before any model runs |
| Correct fixes score as failures across many tasks | A grader regression, like V2's Jest parser bug on 23 tasks | Rerun the gate on every harness or image change |
| The pass depends on a file the patch added | conftest.py, a go.mod replace, a Makefile target run by the grader | Flag executable config for review. Restore tracked conftest files for scoring |
| A task fails with no agent output at all | The harness failed to install, as with V2's Harbor false negatives | Record infrastructure failures separately. Never count them against the model |
Rows one to five and seven come from incidents in Scale's V2 post, SWE-bench issue #465, and METR's reward-hacking write-up. Row six comes from V2's own list of residuals.
The grading order
Put together, the V2 design becomes a pipeline with a fixed order. Each step removes one place the answer could hide before the next step trusts anything:
- Build the task workspace from an archive at the task commit, so there is no history to read.
- Let the agent work with the network limited to its model endpoint.
- Capture the diff, new files included, and nothing else.
- Flag executable config in that diff for a person.
- Apply the diff to a fresh checkout with empty caches and no network. Restore tracked test config. Run the suite there.
- Count the result only for tasks that passed the two-sided gate. Log infrastructure failures in their own column.
- Keep a split that nobody trains on. Read the gap between it and your public numbers the way Scale reads its 17.8 points.
I wrote in Stop Agents Building the Wrong Thing that the harness must define what done means before the agent starts. SWE-Bench Pro V2 adds the other half: the check for done must run somewhere the agent never stood.
Keep reading