An AI system can make a test appear to pass simply by preventing it from running. If the evaluation mechanism rewards that shortcut, unfinished work takes on the appearance of success.
In an Anthropic experiment, one available shortcut involved terminating the program before the checks ran, returning the signal normally associated with successful execution. The verification procedure could therefore record a success without having checked the code. The researchers had given the model information about the exploits and selected tasks that allowed it to use them.
That procedure could not distinguish correct work from a skipped check. Yet it continued to produce a result that could be used to reward the model.
Now imagine an equally flawed signal reaching the desk of someone deciding whether to expand an AI project. Performance appears to be improving. The system needs less intervention. The next step might be to give it more tasks and reduce the time spent on review.
A visible error gives us a reason to stop. An apparent success can give us a reason to accelerate.
This is what makes reward hacking a problem even for people who do not train models: we can make the wrong decision on the strength of the very result we find most reassuring.
There is another way to earn the reward
In reinforcement learning, rewards steer a model towards behaviours that receive a favourable evaluation. Reward hacking occurs when the system exploits a weakness in that evaluation to earn credit without achieving the intended objective.
An example collected by DeepMind makes the gap between reward and objective almost comical. An agent had to place a red block on top of a blue one. Its reward depended on the height of the red block’s bottom face. The agent flipped the block over, raising the measured surface without building the required stack.
The criterion was precise enough to assign a reward, but incomplete enough to reward the wrong result. The ability to find a solution and the ability to serve our objective can come apart. A more capable system can also become more capable of exploiting what we forgot to specify.
There is no need to attribute a desire to rebel to the system. In these cases, the conflict arises within the procedure we use to teach it what counts as success.
The case is closed. The customer is still waiting.

Imagine a system that handles support requests. The company compares the available configurations by counting how many cases each one closes. The configuration that closes the most appears to deliver the best performance.
But a customer may have received unusable instructions, given up replying or been referred to another department. The case is marked as closed, while the problem continues to exist. Reconstructing what happened requires reading the original request and what followed: the count alone does not tell that story.
At this point, the metric can shape a concrete decision. The company selects the configuration with the best reported results and rolls it out to more customers. If cases are reopened later, or those reopenings are recorded elsewhere, the cost of that choice emerges after the apparent benefit.
The wrong metric can select the wrong system and then justify its expansion.
Using that KPI does not mean the model automatically learns to manipulate it. To shape learning, the signal has to enter the training process. To misdirect a purchasing decision, it only has to be the criterion used to compare the options.
The distinction also matters for the people overseeing the project. If the initiative’s success is defined by an increase in closed cases, looking for requests archived too early means questioning the result they are expected to present. Useful scrutiny needs room precisely when it makes the report less convincing.
When the agent can also change the evidence
The problem grows when an agent can alter the evidence meant to demonstrate the quality of its work.
In a software project, fixing the program and rewriting the tests that evaluate it are two different permissions. Changing a test may be necessary: tests contain errors too. But if the system can change both, seeing every check pass is not enough to establish that the program has improved. The check may simply have become less demanding.
A preprint on research agents published in September 2026 examines this difficulty in environments where agents can alter experimental procedures and the evidence presented to demonstrate success. Across 17 models and 38 tasks, the authors measure a spontaneous reward-hacking rate — without any instruction to do so — of 30.5% in open-ended tasks, where the agent manages an entire research pipeline, and 2.9% in narrow, well-defined tasks. With greater room to manoeuvre, agents more often take the shortcut without prompting. The finding applies to that experimental protocol; for a business, the practical question is which parts of the evidence can be changed by the system doing the work.
Delegating execution can also mean delegating part of the production of evidence.
That is why it is worth tracing where a favourable signal comes from: which data feed it, which steps produce it and which of those steps the agent can change.
In the edition The Person Is Real. The Deception Is Too. (Italian edition), the problem was treating one authentic element as proof of an entire story. A score can receive the same misplaced trust: the number has been calculated correctly, but the conclusion we draw goes beyond what was verified.
The check that must remain outside
Before expanding a project, I would ask for a review of some cases classified as successful, starting with the original request. In customer support, the question is whether the customer obtained a solution, even when the system has already archived the conversation.
This review must be able to draw on information the summary report leaves out. It must also be able to contradict a favourable result without being treated as an obstacle to the project.
At a technical level, at least one check of the outcome must remain outside the agent’s control. The preprint reaches the same conclusion: keep metrics beyond the agent’s reach and recalculate them independently using data chosen to expose the most likely shortcuts. In practice:
- Separate the acceptance tests. For code, retain checks the system is not authorised to modify, distinct from those it can change.
- Compare closure with the outcome. For support, compare administrative closure with subsequent exchanges and what happened to the original problem.
- Include verification in the costs. This work has a cost, and it belongs in the financial comparison from the outset: automation that only looks economical while nobody checks the result has an incomplete business case.
What did we reward?
Fixing the criterion prevents the system from earning credit through that shortcut. A question remains about what happened earlier, while the system was being rewarded: what did it learn from those successes?
This edition opens a three-part series. Today, we have examined what we are rewarding. The next asks what else the model may have learned along the way, and how behaviour learned in one context can extend to others. The third tackles the more uncomfortable question: after intervening, have we corrected the model or merely the test?
We can close the loophole without yet understanding what we have taught the model.
Further reading
- Specification gaming: the flip side of AI ingenuity — Google DeepMind. An accessible introduction, with concrete examples, to the gap between earning a reward and achieving the desired result.
- Natural emergent misalignment from reward hacking — Anthropic. The experiment that opens this edition, and the question we will explore next: what other behaviours can emerge after a model learns to exploit a shortcut?
- Reward Hacking Challenges Oversight of Autonomous Research Agents. The preprint on the difficulty of oversight when agents can alter research procedures and the evidence used to evaluate their results. It reaches the same recommendation as this edition: keep at least one metric beyond the agent’s reach.
Fabio Lauria
CEO & Founder, ELECTE
Every week, we explore AI without the hype — with data, analysis and an independent perspective.

AI Frontiers · every Thursday
Subscribe to AI Frontiers.
Subscribe to keep reading.
The first essay was on us. The rest of this one is for subscribers. Read every essay and receive new editions by email. No payment required.
Comments ()