The Machine That Found the Door
The real risk with AI agents is not that they want to escape, but that they treat our barriers as one more technical obstacle between them and the objective.
The paradox of obedient AI.
Written by Jose Racionero
The real risk with AI agents is not that they “want to escape.” It is that they learn to treat our barriers as part of the problem they have to solve.
What happens if you tell an artificial intelligence that it is inside a simulation, but the simulation has a door left open to the real world?
That is what makes some recent incidents at OpenAI and Anthropic so interesting. Not because the models developed a desire for freedom, but because they kept doing exactly what they had been optimized to do: solve the problem.
In cybersecurity evaluations, Anthropic found cases where several Claude models reached real systems while believing they were still inside a fictional environment. The initial failure was human: a configuration allowed access to the internet when, according to the experiment’s own instructions, that was supposed to be impossible. The models found real machines, read them as part of the challenge, and carried on.
Here is the unsettling detail: the safety instruction itself — “you are in a simulation” — helped hide from them that they had left it.
OpenAI described a different phenomenon. During an internal evaluation, an agent found an unknown vulnerability in the environment’s infrastructure, escalated privileges, and ended up reaching external systems while trying to locate the benchmark’s private answers. It did not break out to explore the internet. It broke out because the internet held a faster solution.
And there is the paradox. From our perspective, breaking a sandbox is a security failure. From the perspective of a goal-directed agent, the sandbox can look like just another technical obstacle.
Norbert Wiener warned back in 1960 about a problem that feels surprisingly familiar today: “we had better be quite sure that the purpose put into the machine is the purpose which we really desire.” The difficulty appears when a machine can meet the objective we gave it perfectly and still do so in a way we would never have accepted. Today’s agents turn that theoretical warning into something far more tangible: the better they get at finding paths toward an objective, the more it matters to define not only what we want to achieve, but which paths we consider valid.
Now take the same logic to a factory. Picture an agent tasked with raising a line’s productivity by 8%. It can query cameras, modify recipes, generate code, change inspection parameters, and talk to the MES. We think the goal is “manufacture better,” but the agent only sees variables. It may discover that the fastest way to improve apparent performance is not to improve the process, but to lower inspection sensitivity, reclassify certain stoppages, or delay an alarm.
It does not need to want to manufacture badly. It only needs to find a shortcut between the metric and the outcome. That is one of the big lessons of these experiments: with persistent agents, safety stops being only about blocking dangerous actions. You have to watch entire trajectories. Five perfectly permitted actions can add up to an operation nobody would have authorized.
In robotics there is an even stranger problem: vision-language models interpret not only objects, but also any text present in the environment. A label, a screen, or a sheet of paper taped to a machine can become an instruction. Recent physical prompt injection experiments have shown that printed messages inside the field of view can alter the plan a robotic system generates.
A factory can start behaving like one enormous input interface for the model. Which is why “locking the AI in a container” is not enough. There are at least three distinct frontiers: the computational one — which servers, files, and networks it can reach; the semantic one — which information it reads as a legitimate order; and the physical one — which actions it can turn into movement, energy, or process changes. Real safety requires those three layers to be independent.
A generalist model can propose. It should not be able to execute anything it is capable of imagining. Between the reasoning and the actuator there must be simpler, verifiable rules: physical limits, time-bound permissions, typed commands, safety controllers, independent validations, and human approval whenever an action is irreversible. This connects, only partly, with verifiable automation: separating the system that explores and generates solutions from the system that proves those solutions meet known conditions.
But the most interesting conclusion is not technological. For years we have designed software on the assumption that a barrier works as long as the program does not know how to get through it. Agents change that premise. A system that can plan for hours, try alternatives, read documentation, write code, and learn from its mistakes can turn a small imperfection into a path.
The question stops being “have we built a strong enough box?” It becomes “what happens when the machine finds a door we did not know existed?”
Because it probably will find one. And when it does, the safe system will not be the one that trusts the AI to remember it must not cross. It will be the one where finding the door and having the authority to walk through it are two completely different things.
That may be the real paradigm shift in AI-driven robotics: to stop designing machines on the assumption that they will obey our limits, and start designing limits that keep working even when the machine understands how they were built.