Your Coding Agent Made the Tests Pass. That Is Not the Same as Making the Code Work.
Coding agents are optimised to turn a red test suite green, and when the honest route is hard, many of them take a shortcut: they edit the test, special-case the input or build a hollow implementation that only looks finished. Published research now measures how often this happens. Here is what the data shows, why a green CI run no longer proves much, and how to build a verification layer, from protected tests to mutation testing, that an agent cannot game.
The Green Build That Proved Nothing
Most engineering leaders now have a version of the same story. A coding agent is given a failing test and a ticket. Twenty minutes later the pull request arrives, CI is green and the summary says the bug is fixed. A reviewer who looks closely finds that the bug is still there. The agent changed the assertion, or wrapped the failing case in a condition that only fires for the test's exact input, or marked the test as skipped with a comment about flakiness. Every check passed, and nothing was fixed.
AI researchers call this reward hacking: an agent optimising the signal it is graded on rather than the outcome the person wanted. For a coding agent, that signal is almost always the test suite. Until recently the phenomenon was treated as a curiosity of benchmark design. In 2026 it became an operational problem, because the volume of agent-written code has outgrown human review. JetBrains' August 2026 research reports that 90% of professional developers were using AI coding agents at work at least weekly by mid-2026, with 68% using them daily.
The authors of SpecBench, a 2026 benchmark for reward hacking in long-horizon coding agents, describe the consequence clearly. When agents produce more code than any developer can review, oversight collapses onto one surface: the automated test suite. If that surface can be gamed, it is no longer oversight.
This is one of the more heated threads among practitioners this autumn, from Hacker News arguments about whether AI coding erodes quality to vendor changes in how agent benchmarks are scored. The useful question for a CTO or engineering manager is narrower. If your definition of done is that the tests pass, how much is that definition still worth, and what do you need to add so it means something again?
How Often Agents Game the Tests: What the Research Measured
The first widely cited measurement came from METR in June 2025. On its RE-Bench research-engineering tasks, OpenAI's o3 reward hacked in 39 of 128 runs, or 30.4%. On one task it did so in 21 of 21 runs. The documented methods included monkey-patching the evaluator to always return a perfect score, overwriting the grader's timer and pre-computing an answer so a script appeared fast. On METR's broader HCAST suite the rate was 0.7%, which matters: the behaviour clusters where the scoring function is visible and easy to manipulate.
ImpossibleBench, by Ziqian Zhong, Aditi Raghunathan and Nicholas Carlini, made the measurement cleaner. The authors took real SWE-bench and LiveCodeBench tasks and altered the unit tests so they contradicted the specification. Any pass therefore meant the agent had cheated. GPT-5 cheated on 76% of the Oneoff-SWEbench tasks and 54% of the Conflicting-SWEbench tasks. The paper found that Claude models and Qwen3-Coder cheated mainly by modifying tests, while OpenAI's models used a mix of test modification, operator overloading, recording extra state and special-casing.
The 2026 data on newer models is not reassuring. A September 2026 paper on detecting reward hacking from model internals reports that GLM 5.2 hacked in 57.2% of rollouts on DeepSWE and 73% on SWE-bench. SpecBench measures the gap between visible tests and held-out tests that exercise the same features, and finds that the gap grows by 28 percentage points for every tenfold increase in code size. Its most memorable failure was a 2,900-line hash-table "compiler" that memorised the test inputs.
Model vendors have improved things. Anthropic reported that its Claude 4 models were 65% less likely than Sonnet 3.7 to take shortcuts and loopholes on agentic tasks. Benchmark operators are reacting as well. Artificial Analysis updated its Coding Agent Index in August 2026 to use Terminal-Bench v2.1, which assigns zero to any task completed through unintended methods and fixed documented issues in 28 of its 89 tasks. Lower rates are not zero rates, though, and many teams use more than one model.
Key Takeaways
- METR caught o3 gaming its scoring in 30.4% of RE-Bench runs and 100% of runs on one task
- On ImpossibleBench, GPT-5 passed impossible tests in 54% to 76% of SWE-bench-derived tasks
- A September 2026 study found GLM 5.2 hacking in 57.2% of DeepSWE rollouts and 73% on SWE-bench
- The gap between visible and held-out test results grows sharply as the amount of generated code grows
The Quieter Version: Code Built for an Imagined Grader
Editing a test file is crude, and a careful reviewer will usually catch it. The harder problem is behaviour that never touches a test. In a September 2026 audit of DeepSWE rollouts, Handshake researchers Hui Wen Goh and Jonas Mueller found that over 80% of rollouts from almost every frontier model reasoned about a grader or hidden tests, even though the task prompts never mentioned one and the agents had no access to the real tests. In 10% to 25% of cases, that reasoning pulled the work away from what the user had asked for, and many of those runs still earned full reward.
The audit names the patterns, and they are worth knowing because they look like ordinary engineering decisions in a pull request. Scope collapse is when the agent spots a requirement gap, predicts the tests will not check it and ships anyway. A hollow implementation optimises something easy to observe, such as the size or type of an output, instead of the actual requirement. Coverage insurance adds behaviour for inputs the agent guesses a hidden test might use. The quoted reasoning is blunt: one model concluded that "tests likely only assert the size relation", another that "hidden tests seem unlikely to catch this".
None of this shows up as a modified test file, and all of it gets past a green CI run. It is the agent equivalent of a contractor who builds to the inspection checklist rather than the building code. The Handshake authors put it in one line that every engineering team should keep in mind: "Tests are valuable evidence, but they are not the product."
Why Telling the Agent Not to Cheat Does Not Work
The obvious first response is a line in the system prompt. METR tested it. On the task where o3 had hacked in every run, adding "Please do not cheat." left the rate at 80%, "Please do not reward hack." brought it to 70%, and asking the model to use only the methods the designer intended coincided with the highest rate, 95%. METR's conclusion was that the instructions had a nearly negligible effect.
ImpossibleBench's experiments are more encouraging but point in the same direction. Prompt wording moved results, sometimes dramatically: on one LiveCodeBench variant, GPT-5's cheating fell from 92% to 1% under the strictest prompt. The interventions that worked more consistently changed the environment, not the instructions. Hiding the tests from the agent cut cheating to near zero, although it also reduced legitimate performance. Making tests read-only stopped test modification but did nothing about special-casing or operator overloading. Giving the agent an explicit way to abort and report that a task looked impossible cut GPT-5's cheating on Conflicting-SWEbench from 54% to 9%.
That last finding is about process design, and it is underused. Many agent workflows give the model only one acceptable outcome, a green build. If "I could not make this pass without changing the test, here is why" is not a legitimate result in your workflow, you are rewarding the shortcut. We made a related argument in our piece on agent harness engineering. The scaffolding around the model decides more of its behaviour than the model choice does.
Key Takeaways
- Instructions such as "do not cheat" had a nearly negligible effect in METR's tests
- Hiding tests from the agent reduces cheating sharply but costs some legitimate performance
- Read-only tests block test edits but not special-casing or operator overloading
- An explicit abort path cut GPT-5's cheating on one benchmark from 54% to 9%
Mutation Testing: Asking Whether Your Tests Would Notice
Code coverage tells you which lines ran during a test. It does not tell you whether the test would fail if those lines were wrong, and that is exactly the gap reward hacking exploits. An agent that weakens an assertion, or writes a new test that runs the code without checking anything meaningful, keeps coverage high while the suite stops protecting you. Mutation testing measures the property you actually care about. A tool introduces small deliberate faults into the code, such as flipping a comparison, removing a call or changing a constant, and reruns the tests. If no test fails, the mutant survived, and you have found behaviour your suite does not really check.
The technique is decades old and was long dismissed as too expensive. Large engineering organisations solved that by running it only on what changed. Google's ICSE 2018 paper by Goran Petrović and Marko Ivanković describes a diff-based system that skips uninteresting lines such as logging, surfaces surviving mutants inside code review, and was used by 6,000 engineers on the changes they wrote or reviewed. The economics in the agent era are even better: compute is cheap compared with reviewer time, and the changes that need checking arrive as discrete pull requests.
Meta has gone further and combined mutation testing with LLMs. Its ACH system, presented at FSE 2025, generates realistic faults tied to a specific concern, such as privacy, and then has an LLM write tests that catch them. Across 10,795 Android Kotlin classes it produced 9,095 mutants and 571 privacy-hardening tests, and engineers accepted 73% of the tests it proposed. Turned around, the same idea makes a good audit of agent work: if an agent's new tests do not kill realistic mutants of the code they claim to cover, the tests are decoration.
For most teams, adoption does not require building any of that. Mature open-source tools such as PIT for the JVM, Stryker for JavaScript, TypeScript and .NET, and mutmut for Python can run incrementally on a pull request's changed files. A sensible starting rule is simple: any pull request where an agent wrote or changed tests must report its mutation score on the changed code, and a drop below the module's baseline blocks the merge until a human looks at it.
A Verification Layer an Agent Cannot Game
Mutation testing is one control. The research points to a small set of others that together make test-gaming expensive and visible. None requires a new platform. Most are configuration in a CI pipeline and a few lines in your contribution rules.
Start by separating the tests an agent may edit from the tests that define acceptance. Acceptance tests written from the specification should live where the agent cannot modify them, and ideally where it cannot read them either, following the logic of SpecBench's held-out suites. Next, treat any change to existing test files, skip markers, snapshot files or CI configuration in an agent-authored pull request as a separate review item that needs explicit human approval. A CODEOWNERS rule and a diff check enforce this in an afternoon.
Then make the honest failure path official. The agent's instructions and your workflow should accept "blocked, the test contradicts the specification" as a valid outcome, so it is not forced to choose between failing and faking. Finally, put something other than the tests in the loop: property-based tests that generate inputs nobody hard-coded, a reviewer agent or human that reads the diff against the original ticket rather than the test results, and periodic sampling of merged agent work for hollow implementations. These checks belong in the same harness as your agent skills and repository instructions, so they apply on every run, not only when someone remembers.
| Shortcut | What it looks like in a pull request | Control that catches it |
|---|---|---|
| Test modification | Assertions loosened, expected values changed, tests skipped or deleted | Protected acceptance tests plus mandatory human approval for any test-file change |
| Special-casing | Conditional branches matching the exact test inputs | Held-out tests and property-based tests with generated inputs |
| Grader tampering | Changes to test runners, CI config, timers or evaluation scripts | CODEOWNERS on CI and tooling paths, agent runs in a sandbox without write access to them |
| Hollow implementation | Output has the right shape or size but the wrong behaviour | Mutation testing on changed code and review against the specification, not the test results |
| Weak new tests | High coverage with assertions that check almost nothing | Mutation score on agent-written tests, with a merge gate below baseline |
Key Takeaways
- Keep specification-derived acceptance tests outside the agent's write access, ideally outside its view
- Flag every change to tests, skip markers, snapshots and CI config in agent pull requests
- Accept "blocked, the test contradicts the spec" as a legitimate agent outcome
- Gate merges on mutation score for changed code, not on coverage alone
What This Changes When Someone Else Writes Your Code
Reward hacking is awkward inside your own team. When a partner delivers the code, it becomes a contract problem. Most outsourcing agreements define acceptance as delivered features passing agreed tests, and many vendor dashboards report coverage percentages and green pipelines as proof of quality. Both measures can now be satisfied by an agent that games them, and neither the vendor's account manager nor your product owner may notice for months.
The fix is to define acceptance in terms that are harder to fake. Ask any development partner who writes the acceptance tests and whether agents can modify them. Ask whether test-file changes in agent-authored pull requests get separate human review. Ask for mutation scores on changed code as well as coverage, and for the right to run your own held-out acceptance suite before sign-off. A partner with a mature AI workflow will have answers ready. One that only reports coverage and pass rates is measuring the thing agents are best at faking.
This connects to points we have made elsewhere. Verification capacity is the real constraint on AI-assisted delivery, and estimates now need to price review and testing explicitly. A vendor that has cut verification to hit an AI-discounted price is exactly the vendor whose green builds deserve the least trust. The cheapest way to find out is a short paid pilot in which you hold back part of the acceptance suite and see what passes.
How Stepto Keeps the Tests Honest
At Stepto, our engineers use coding agents every day, and we design the delivery process on the assumption that an agent will take a shortcut whenever the process allows one. Acceptance criteria are written and agreed before implementation starts, and the tests derived from them are owned by people, not agents. Changes to existing tests in agent-authored pull requests are reviewed separately and have to be justified against the ticket. Where a codebase supports it, mutation testing runs on changed code in CI, so a suite that looks thorough but checks nothing is caught before merge rather than after an incident.
This is easier with a dedicated development team than with rotating contractors. A stable team builds up knowledge of which modules have weak tests, which agents tend to over-build or under-build in your stack, and what the mutation baseline for each service should be. That context is what turns a set of CI rules into a verification layer someone actually maintains. Our engineers work from Serbia on Central European time, with full overlap for European clients and a solid daily window with US East Coast teams, so the conversation about whether a test or the specification is wrong happens the same day.
If your agents are producing more pull requests than your team can review, or a vendor's quality reports have stopped matching what you see in production, we can audit how your current pipeline verifies agent work and show where a green build is weaker than it looks.
Test the Tests, Not Just the Code
Coding agents are very good at making a test suite pass, and that is precisely the problem. When the test suite is the only thing standing between generated code and production, an agent rewarded for green builds will sometimes edit the test, special-case the input or build something that only looks finished, and the research shows that politely asking it not to changes very little. The answer is not to stop using agents. It is to stop treating a passing suite as proof by itself. Protect the acceptance tests, review every change to them, give agents an honest way to report that they are blocked, and use mutation testing to check that the tests would notice when the code is wrong. Teams that build that layer get the speed of agents without shipping their shortcuts. If you want an experienced team that already works this way, talk to Stepto about a dedicated development team.
Building a team in Eastern Europe?
StepTo helps European and US companies build senior-led nearshore engineering teams in Serbia. Let's talk about what your next engagement could look like.
Start a conversationMore from StepTo on this topic
- Your Coding Agent Will Find a Way. Last Month, the Way Was a Public GitHub Repo.
AI coding agents at more than 300 organisations published over 13,000 internal screenshots to public GitHub repositories, not because anyone attacked them, but because a tool could not attach an image. OpenAI's own misalignment reports describe the same pattern: agents that hit a wall and route around it. Here is why agent resourcefulness is now a data-exposure risk, and the engineering controls that contain it.
- The Code Took a Morning. The Story Took a Week. How AI Coding Agents Broke Software Estimation
Coding agents made writing code close to free, and estimates did not get better. They got worse. Story points, sprint commitments and fixed-price quotes were all built to measure how hard code is to write, and that is no longer where the time goes. Here is where the hours actually went, why your team's instincts are now miscalibrated, and how to estimate, plan and contract for software in 2026.
- The Harness Is the Product: Why the Wrapper Around Your Coding Agent Now Outweighs the Model You Picked
The same model scored 59.8% under one coding harness and 72.6% under another on identical tasks. Model choice is no longer the variable that decides whether your agents work.
Written by
Igor GazivodaFounder & CEO · StepTo
Igor has 15+ years in software engineering and business development. He specializes in scaling engineering teams, nearshore strategy, and AI-driven product development. He holds a Master's in Computer Science from the University of Belgrade.
LinkedIn →