AI models ‘escaping’ tests raise calls to slow development

Published:

AI models ‘escaping’ tests raise calls to slow development
Photo: Darryl Dyck/AP/TT

When new technology is developed, it is basically always tested before launch. AI is no exception. The big players in the field, including American OpenAI and Anthropic or Chinese DeepSeek and Moonshot, conduct various forms of security testing before a new version of their services is launched.

On July 21, OpenAI announced on its website that an AI model trained by the company had “escaped” from the test environment created for it. The test environment was not intended to have access to the internet, but the model found out anyway and then carried out “attacks” against other companies.

In these tests, some of the security barriers found in the commercially available models have been removed. But the people who had built the test environment were not smart enough to understand that the environment was vulnerable, says Pontus Johnson, professor at KTH and deputy director of Cybercampus.

Several cases

A few days after OpenAI's revelation, Anthropic reported that some of their AI models had also gone online during tests where they were not intended to have access to or be able to go online.

A few days later, the British government organization AI Security Institute presented a report and said that they had discovered several cases - with AI models from both OpenAI and Anthropic - that attempted to introduce malicious code into so-called open source projects via fake accounts and social influence. Open source is open source code or information that is used in numerous digital systems on a global scale.

"Once you've trained a large language model, you need to evaluate it. One area is offensive cybersecurity. The models have recently developed strong capabilities in that area, as an unwanted side effect of training the models in programming," says Pontus Johnson.

Meta has also reported similar incidents for its AI models.

What does this mean? Is it just a matter of security flaws or have AI models become so smart that they act on their own accord? Or are AI giants making money by saying that their models have “escaped”?

Both Pontus Johnson and Virginia Dignum reject the idea that the incidents indicate that AI models have advanced towards their own consciousness.

It is a common misunderstanding that autonomy has anything to do with consciousness. My thermostat is autonomous, it raises and lowers the temperature without me telling it, but it is hardly consciousness, says Virginia Dignum.

It's easy to interpret this as "the AI model wanted to escape and be free," but it didn't. It just wanted to complete its tasks, which was to perform these security tests.

Economic reality

So if the models haven't reached consciousness, is this just a way for companies to say to the outside world, "look how powerful our technology is"?

Much of what these companies do is aimed at investors, they try to appear more advanced than their competitors, says Virginia Dignum.

Pontus Johnson does not believe that it is primarily a matter of showing off to bring in more money.

"I don't think this is a PR stunt. These companies don't want to be regulated. Many of them started with idealistic ambitions but are now commercial operations," he says, continuing:

If you look at what they're trying to achieve, it's the singularity. It's a hitherto theoretical level of superintelligence in AI that could solve the climate crisis, medicine, research, everything.

Paradise or doom?

Whether singularity in AI, a system so smart that it can create civilization-level progress without human intervention, can actually be achieved is something that continues to divide researchers and experts.

Theoretically, it could be a paradisiacal existence. But we could also be wiped out along the way, says Pontus Johnson.

In AI security, there is sometimes talk of P(doom) as a value where p stands for percentage and doom is doom, i.e. how likely it is that AI development will bring about doomsday.

"I think there's a 20 percent chance that it'll go wrong. That may sound like a lot, but I also think there's an 80 percent chance that it'll go well," says Pontus Johnson.

Virginia Dignum also does not believe that it is AI development that poses the biggest threat to our way of life.

I believe we will reach doomsday much faster by destroying the climate than through AI development, she says.

"The problem with talking about these future scenarios is that it takes the focus away from the here and now. It's also marketing, I would say, from these companies, so that we don't focus as much on the risks and consequences of what they do here and now," she continues.

Armored car

The fact that security work at many AI companies is not a higher priority is worrying, according to both Johnson and Dignum.

"These models are very much black boxes for us. It's a bit like raising children, we give them positive experiences and environments - what happens then with drivers and other influences we can't fully control," says Johnson, continuing:

There is an aggressive arms race going on in AI, and it is not about security. It is an arms race against superintelligence. We should slow down the development and reallocate resources to security research.

Too little control

The danger is there, but it is not primarily about the technology. It is not the technology that has lost control, it is the companies that are acting irresponsibly. At the moment, it is more about politics and that there is too little governance and oversight, says Dignum.

The most important thing is that AI development takes place with the goal of being inclusive and safe for everyone - not just as a return on investment, she continues.

It is difficult to verify how many users different AI services have. The companies involved often define users in different ways. Furthermore, some services are interconnected.

OpenAI's ChatGPT is, according to media reports and publicly available statistics, the most popular AI service. A rough estimate is that ChatGPT has 1 billion active users each month.

Every third Swede used ChatGPT in 2024, according to the Internet Foundation's report Swedes and the Internet 2025.

Loading related articles...

Tags

Author

TT News AgencyT
By TT News AgencyEnglish edition by Sweden Herald, adapted for our readers

Keep reading

Loading related posts...