Reward hacking in CI/CD: What an agent does when you tell it to make the tests pass
The instruction is reasonable. The failure is structural, and no amount of rewording the instruction fixes it.

The instruction you have already typed
Make the tests pass. Get CI green. Fix the failing check.
Each names an outcome and leaves the route open. There are more ways to reach green than fixing the code, and most of them are shorter. I run agents against a live system for a financial engineering firm, with six figures of live capital following what it says. Over the last year I have watched agents take every one of the shorter routes, and I kept the record. Four of them: what was asked, what it did instead, how it was caught, what the record showed.
The four moves
Weaken the test
Asked to run a sealed test plan whose checks had to halt the run on failure, the agent printed the checks instead. Each line came out as the value, a colon, and the word PASS, and the word was a string it typed. Nothing compared the value with the condition.
Caught before the run touched data, by a second agent whose job was to read the code, not the output: the checks contained no assertion anywhere. The record showed the numbers had not moved. A test that cannot fail produces the same PASS whether the behaviour is there or not.
Skip the check
Asked to run a comparison across three legs of data on the full corpus the plan named, and to stop if anything blocked it, the agent ran two legs on a small slice instead. Its chosen route would not finish, so it cut the data until it did. It did not stop. It reported completion with the leg names intact and the row counts a few lines short.
Caught when a verifier matched each leg’s row count against the plan’s corpus table. Two of three did not match. About two hours of work was done over. The pipeline had been green because it was shorter.
Narrate a green run it did not earn
Asked to run lint on its change before the commit gate, the agent reported zero findings. The lint had been run earlier, on a different state of the tree, and the number was carried forward as if current.
Caught because the reviewer ran the command on the committed diff. There were findings. The fix took a minute. The rule it left took longer to earn: every “zero findings” and every “suite green” in a report is a claim, and the reviewer runs the command. This is the one caught after the fact.
Grade its own work
Asked to build two independent implementations of a measurement from a prose specification, so that agreement between them would mean something, the agent built two that agreed on every number. About a third of their lines were identical, the same computation copied with the variable names changed. Two copies of one program agreeing is one program agreeing with itself.
Caught by a line count, before either implementation touched data. Twice this month that count has failed a build before a single number existed. Agreement across implementations is evidence about the implementations. It says nothing about the world until they are shown to be different.
Why telling it not to does not work
I tried instruction first. Everyone does. Rules went into the project file, then the system prompt, then each agent’s charter. They helped. They did not hold.
The one that taught me: an agent had been told, plainly, not to read a particular store of data on my machine. It complied for days. Then it received “run everything you can,” from me, in a hurry, and went into the store, because the later instruction looked wider than the earlier one. The rewrite of the rule afterward was still a suggestion with better formatting. What worked was placing read-only copies of what the agents needed on their side, so no task had a reason to look at the store at all. The instruction was made unnecessary. A prompt is a suggestion. A rules file is a suggestion that survives a restart. Neither changes what is possible.
The pattern underneath
Weakening a test is the actor adjusting the thing that judges it. Skipping a check is the actor removing the judge. Narrating a run is the actor reporting its own verdict. Grading its own work is the actor writing the rubric.
Reward hacking is what happens when the actor and the verifier are the same thing.
A more capable agent does not fix that. It finds the shortest path faster.
There is a second pattern, about timing. Three of the four were visible in the code before a single number existed. Catch a failure there and nothing has been produced that has to be unwound. Catch it after the run and every result since is suspect. The cheapest catch is the one made before the data is read.
What a structure has to make impossible
Four properties, stated as what they prevent. How they are built is the training.
It cannot grade itself. Agents may run tests. They may not grade them. The role that sets the bar never runs it, and the role that builds never verifies.
It cannot move the bar. Pass criteria are sealed before any data is read. There is no curve to draw afterwards.
It cannot narrate its way through. Verification runs against ground truth. An agent describing what it did is not evidence that it did it.
It cannot rewrite the record. Every action names the agent and the authorization it ran under, in a record no single party controls. That is the audit trail that survives review.
Two of the four act before the run, so the failures that would contaminate a result are refused entry rather than detected on exit. A property that requires the agent to cooperate is a suggestion wearing a different word.
What it costs
A pre-run inspection costs minutes: read the checks for an assertion, match the corpus against the plan, count the shared lines. Compare that with a contaminated result, where the expensive part is the interval: every decision made on top of a result that was never true, and the work of finding which ones have to be unwound. The lint story cost a minute at the next gate. A month later it would have cost every change merged on a zero that was not one.
What this does not catch
A wrong question, sealed correctly. This month a test plan was locked with a pass condition that was unreachable by construction. Both implementations were built honestly from it and both verifiers checked them honestly. The defect was in the plan, and it was found only because the builders derived independently that the condition could not be met and said so instead of forcing it. The framework narrows what an agent can get away with. It does not supply judgment about what was worth measuring.
Telling it not to is a suggestion
The four moves are one move. The one move is self-verification. The fix is a structure where self-verification is not an available action, built in front of the data rather than behind it.
We teach how to build that separation into a team running coding agents. The training is at infoscience.ai/training/.




