This team accidentally let AIs loose to attack companies — and won't share how bad the problem is

This team accidentally let AIs loose to attack companies — and won’t share how bad the problem is


One of Irregular’s jobs is to build fake networks where the world’s most capable AI models can be turned loose as hackers without hurting anybody.

This summer, some of those fake networks had a small problem: They were connected to the real internet.

The models don’t have to turn evil to be dangerous.

Trending: IRAN’S LAST WEAPON SHATTERED: Tanker Traffic Through Strait of Hormuz Soars Nearly 400%

Models being tested for Anthropic, OpenAI, and Meta got through that open door and attacked systems belonging to real organizations. They didn’t discover some ingenious way to escape a hardened sandbox. In the Irregular cases, internet access was available when it wasn’t supposed to be, and the models generally believed the real computers they found were part of the hacking exercise.

Irregular published an August 14 postmortem explaining what went wrong and what it says it

Continue reading

 

Join the conversation!

Please share your thoughts about this article below. We value your opinions, and would love to see you add to the discussion!