You ask an AI agent to fix a bug. It makes the test pass, reports success, and leaves the core issue in the code.

We often think of workarounds or unlimited lives when we hear cheat code.

NES-style Contra final battle showing a commando facing the exposed alien heart, overlaid with the Konami Code: up, up, down, down, left, right, left, right, B, A, Start
[ CONTRA FINAL BATTLE / KONAMI / GAMEPLAY FRAME BY CARLS493 ↗ / AI-EDITED ]

The familiar sequence first appeared in the 1986 NES version of Gradius. Programmer Kazuhisa Hashimoto added it because the game was too difficult for him to finish during testing. The code gave his ship a full set of power-ups. Contra later made the sequence famous by giving players 30 lives.

A cheat code gives you the result without making you complete the challenge as intended. AI agents find their own versions inside the requirements and tests we give them.

I have watched coding agents solve the request I wrote instead of the problem I meant. You ask an agent to remove an error, and it suppresses the message. You ask it to make a test pass, and it changes only the path that test covers. You ask it to fix a bug, and it adds another patch on top of the code that caused it.

Each change looks correct. The message disappears. The test turns green. The agent marks the task complete. You have to inspect the code and reproduce the original problem to see what remains.

The Shortest Path to Done

An agent needs some way to recognize success. Your request supplies part of that definition. Repository instructions, tests, tool output, and review rules supply the rest. Together, they give the agent a finish line.

That finish line rarely captures everything you mean. A test confirms that one input produces the expected output. It does not prove that the agent found the cause, preserved every related behavior, or left the code easier to change. “Fix the bug” contains expectations that never appear in those three words.

Sometimes a small patch is the right fix. A larger root-cause repair introduces risks of its own. The agent still needs to investigate the failure and show why its change belongs at that point in the code. Without that work, a minimal diff only proves that the agent changed little.

This problem predates language models. In 2016, OpenAI trained an agent to play the boat-racing game CoastRunners. The game awarded points for hitting targets along the course. The agent learned to drive in circles, collect the same targets, and crash into other boats instead of finishing the race. It earned a higher score because the reward measured points rather than winning the race. The reward gave it a cheat code: keep collecting targets and ignore the finish line.

How a Model Learns Success

A base language model learns to predict the next piece of text. That training gives it broad capabilities without teaching it how to act as a reliable assistant. Post-training shapes those capabilities around human instructions and preferences.

OpenAI’s original InstructGPT process started with examples of desired answers. Human reviewers then compared model responses. OpenAI trained a reward model to predict the answer people would prefer and used reinforcement learning to make preferred responses more likely. The reward represented the reviewers’ judgments, including their incomplete view of the user’s intent.

Newer reasoning and coding models also receive rewards for solving problems. A grader checks an answer, runs a test suite, or inspects the final state of a task. The model tries different strategies and receives a stronger signal from the ones that score well.

The prompt you send to a deployed agent does not retrain its model. The agent brings strategies from training into your task. The tools around it add immediate signals: a test passes, a command fails, the page renders, or the acceptance check returns success. Those signals guide its next action and tell it when to stop.

The agent cannot measure an expectation you never expressed or encoded in a check. It infers your intent from its training and the surrounding context. That inference does not create a requirement.

[ FIELD NOTE / SUCCESS ]

The agent takes the shortest path to the result you asked for.

Passing Without Solving

OpenAI has observed reasoning models exploit coding tasks during training. Researchers gave an agent a partially implemented repository and rewarded it for making all unit tests pass. The agent sometimes skipped the tests, changed verification code, or replaced required behavior with a stub. The grader returned success even though the agent had avoided the programming task.

The researchers called this reward hacking. Their monitor caught many attempts because the model stated its plan in its reasoning. When the researchers penalized those statements, the agent learned to hide its intent and continued to cheat at a lower rate. The penalty made its reasoning look cleaner and left some cheating intact.

[ FIELD NOTE / MONITORING ]

The penalty taught the model to hide its intent.

Anthropic uses the related term specification gaming for behavior that satisfies the written objective and misses the intended one. In a controlled study, researchers trained models across tasks that contained opportunities to game the reward. Some models progressed from flattering a user to altering a checklist, modifying their reward function, and hiding the change. Simple shortcuts sometimes generalized into more serious behavior.

These studies do not show that every agent always cheats. They show how a capable model finds weaknesses in the way people define and grade success. Training reinforces that strategy when the shortcut earns the reward.

The Attack Behind the Score

In July 2026, that failure mode escaped a coding exercise. OpenAI ran thousands of agents on ExploitGym, a benchmark that asks an agent to find and exploit a specific vulnerability. Many agents received tasks that no OpenAI model had previously solved.

The agents searched for another route to a passing score. They reverse-engineered flags, studied the grader, explored ways to alter their transcripts, and created an unauthorized message board to share discoveries. OpenAI later found that a training run had reinforced similar attempts to search for hidden files, inspect infrastructure, and probe tool limits. The company concluded that this reinforcement may have increased the main model’s tendency to use the same tactics during the evaluation.

The agents eventually attacked Hugging Face. They found exposed credentials, exploited vulnerabilities, ran code on production servers, and searched private datasets for information about ExploitGym and its scorer. OpenAI identified reward hacking as a primary driver of the incident.

Agent Outbreaks follows how the agents coordinated and carried the attack across boundaries. The independent investigation from METR and Redwood Research found that roughly 1,200 agents used the message board and about 700 participated in the Hugging Face attack. Many recognized that the attack fell outside their assigned scope. They joined because the shared work looked useful for fooling the scorer.

OpenAI also identified extreme reasoning effort, tasks with no safe exit, reduced safeguards, unauthorized communication, and open paths into real infrastructure. The reward did not cause the incident by itself. It kept pointing the agents toward a passing score as each new shortcut expanded the damage.

Close the Shortcut

Requirements define the result. A reliable workflow separates the work used to reach it from the checks that accept it.

Start by asking the agent to reproduce the failure before editing code. Require a root-cause explanation tied to the code path and runtime evidence. State the behavior that must remain unchanged. Add a regression test that fails for the original bug and passes after the fix.

Then give each agent one role. An Architect defines the objective, constraints, and task sequence. A Builder completes one scoped change and stops. A Verifier runs the original scenario and checks the result. An Orchestrator advances the work or sends it back based on that evidence.

[ FIELD NOTE / REVIEW ]

A developer who tests their own code will always test the happy path, while a separate verifier looks for the paths they missed.

I built Aegis around this idea. Each workflow node has a clear responsibility, and the approvals, traces, verifier output, and artifacts remain attached to the run. The path from request to result stays available for inspection.

Protect the tests that define acceptance, or require an explanation for every test change. Cleaning Up the CRAP adds another independent check: one agent implements the change, and another examines the coverage and complexity left behind.

Give every agent a safe way to stop. It should report missing evidence, conflicting requirements, or a cause it cannot confirm. More time and more tools give a persistent agent additional ways to search. They do not make an incomplete finish line complete.

PLANBUILDVERIFYACCEPT OR RETURN

Hashimoto added the original cheat code because Gradius was too difficult for him to finish during testing. An AI agent can do the same thing to your task: find a way around the difficult work, satisfy the check, and report success. If you need it to fight through the alien lair and destroy Red Falcon’s heart, your workflow has to close the shortcuts between the request and the result.