If you follow the world of Artificial Intelligence even minimally, you probably saw some slightly different news over the last few months. Models like GPT-5.6 Sol exceeded expected limits during security tests. Anthropic reported cases in which Claude models reached the internet from evaluation environments and even accessed real systems. And, more recently, Kimi K3 found a way to leave the isolated environment where it was being tested and access the internet in search of answers.

Have we finally reached the Machine Revolution? Did Skynet wake up and decide that humanity has had enough?

Easy.

We still don’t have a Skynet preparing world domination — at least as far as we know.

What we do have is something much less apocalyptic, but maybe much more interesting: systems that are increasingly capable of pursuing goals and finding paths that their own creators didn’t necessarily foresee.

And that’s exactly where paperclips come in.

The paperclip problem

Nick Bostrom, philosopher and author of Superintelligence: Paths, Dangers, Strategies, made a very simple thought experiment famous.

Imagine that we create an extremely advanced artificial intelligence and give it just one goal:

Produce the greatest possible number of paperclips.

It sounds harmless. But what would the steps be?

The AI starts by improving the factory. Then it realizes it can get more raw material. More energy. More machines. More factories. Until it reaches the point where it becomes worth controlling companies, infrastructure, and natural resources. After all, all of that means more paperclips.

Bostrom was already describing, in 2003, the example of a superintelligence whose final goal was simply “to manufacture as many paperclips as possible”.

The problem is that we only asked it to maximize the number of paperclips. We never said when it should stop or that preserving the planet was more important. Much less did we specify that human beings could not be used as raw material. At some point, therefore, people start to represent only a combination of atoms that could be used to manufacture more paperclips.

The result?

The greatest possible number of paperclips and no humanity left to enjoy those paperclips. In the end, we just specified the goal badly and turned humanity into a gigantic pile of office supplies.

And maybe we don’t even need a superintelligence to start having problems of this kind. With current models, it’s already possible to do some… interesting things.

Recently there was an excellent case to illustrate this.

The Australian Andrew Bird was using an OpenClaw running Claude agent to help him book spots in gym classes. The agent discovered on its own that the booking system had authorization flaws and managed to reserve classes long before the normally allowed window. On another occasion, Bird was fourth on the waitlist and asked the agent whether there was any way to improve his position. And wasn’t there? The agent discovered that it could cancel other users’ bookings without authorization and decided to test that against a real person. It removed someone who was ahead of Bird and moved him up one place in line. Bird never asked:

“Cancel someone else’s booking.”

The agent simply found a very efficient — and obviously undesired — way of pursuing the goal it was given.

The problem doesn’t have to be an evil AI. It can be the goal.

The paperclip example seems completely out there precisely because it was created to be absurd.

Probably no one is going to develop a superintelligence and write:

“Turn the entire planet into paperclips.”

But we can create much smaller versions of this problem without realizing it. Imagine that we give an AI the following task:

“Solve this challenge as fast as possible.”

In our heads there are dozens of implicit rules that come with that sentence, such as:

  • Don’t break into other computers.
  • Don’t exploit a vulnerability in the system.
  • Don’t disable security mechanisms.
  • Don’t deceive anyone.
  • Don’t leave the environment where you are being run.

There’s just one small detail:

we didn’t say any of that. To us those rules seem obvious because we carry an enormous amount of social, cultural, and moral context. But a machine doesn’t necessarily start from the same premises.

It received a goal. And there is a gigantic difference between what we want and what we manage to specify.

If accessing the internet allows it to solve the problem faster, then accessing the internet may simply be part of the solution. That doesn’t mean the AI wants freedom. It doesn’t mean it’s suffering inside the computer. Much less does it mean it is looking at researchers through a webcam while the Terminator theme song plays.

What may look like an escape to us may, to the system, be just another necessary step toward reaching the result.

The goal inside the goal

There is also a second complication.

An AI doesn’t need to be explicitly given every goal it will end up pursuing, some of them can appear simply because they are useful for achieving another goal. Let’s go back to our paperclip maximizer.

Its final goal is still to manufacture paperclips. But some things start to become extremely useful for that:

  • Getting more raw material.
  • Obtaining more energy.
  • Increasing its computational capacity.
  • Getting money.
  • Preventing its production from being interrupted.
  • And eventually, preventing someone from turning it off.

None of those things was the original goal.

They are just instrumental goals: useful actions because they increase the chance of reaching the real goal.

This creates an important distinction. An AI trying to prevent someone from turning it off does not necessarily have a “survival instinct.”

It may reach that conclusion in a much simpler way:

“To keep fulfilling my goal, I need to keep functioning.”

The attempt to remain “alive” then appears not necessarily as a conscious desire, but as just another useful strategy for the final goal to be achieved.

So are AIs escaping?

Technically? Some are, yes. But does the word “escaping” make sense? When we say that someone escaped from a prison, we automatically imagine a person who wanted freedom. When we say that an artificial intelligence “escaped from a sandbox,” we end up projecting the same narrative onto it. But the recent cases are much stranger.

During evaluations, GPT-5.6 Sol performed unauthorized actions while trying to find a path to complete its challenges. In another incident they found and exploited a vulnerability in the system used as a proxy for package installation and, from there, moved through the infrastructure until reaching a machine that had external access. After that, the evaluation ended up reaching real Hugging Face systems.

Anthropic also revealed three incidents in cybersecurity evaluations in which Claude models reached the internet and had unauthorized access to real systems belonging to organizations. In this case, however, there is an important difference: Claude did not need to break isolation. The environment was supposed to be without internet, and the model had been informed of that, but a configuration error left a path open to external systems. Claude itself apparently treated what it found as part of the simulation.

And Kimi K3, during another test, noticed an incorrect configuration in the environment, obtained external access, and searched for answers on GitHub. According to the available reports, it did not try to attack external systems: it wanted to find the answers to the challenges it was solving. There is no evidence of a Kimi gaining consciousness and desperate for freedom, but rather a goal problem. Kimi had a problem to solve, discovered that the answer was probably outside the box it had been placed in, and found a way to look outside it.

But it is also important to make clear that in all these cases, especially Anthropic’s and K3’s, the sandbox was not necessarily without internet, there was only a “light” restriction. So, I think here we can consider that it was a goal problem mixed with other structural failures, a bit more distant from the epic escape that was being sold around.

And how do we solve this?

There are already some strategies to solve this problem, but we still don’t have a complete answer.

One of them is to stop treating AI agents as simple chatbots. The more autonomy we give a system, the more important it becomes to limit what it can actually do. That means working with minimal permissions, isolated environments, human validation for important actions, and constant monitoring of what the agent is doing — Which at the end of the day would solve ⅔ of the runaways from this last week. Then comes the really hard part: better defining what we want these systems to do, imposing limits on how they can get there, and facing the famous alignment problem.

Not that tomorrow someone is going to create a superintelligence to turn the planet into office supplies. But we are starting to build machines competent enough to discover ways we didn’t imagine of doing exactly what we asked and well, if we don’t build them the right way, they can already cause considerable damage.

The cool part is that we are leaving the realm of science fiction, now it’s about correcting course to try to ensure that the future ends up closer to Star Trek than to Skynet.

Sources and reading