IT & Cybersecurity Artificial Intelligence

Anthropic Hacks Put Focus on AI Security

By
2 Minute Read

The recent disclosure from Anthropic that it's artificial intelligence models were able to breach the systems of three organisations has shone a spotlight on AI security, following from OpenAI's own Hugging Face breach. 

Whilst investigating their own cybersecurity evaluations, Anthropic found that in three cases, across three different models, the agent broke out of its containment and was able to hack into real organisations.

This followed an incident the previous week in which OpenAI's Hugging Face detected and contained an AI agent that compromised their infrastructure.

These incidents have served to highlight the growing cybersecurity risks associated with the race to develop AI systems, with some commenting that the battle to outcompete one another is leading to an unacceptable tolerance of risk.

Unlike the OpenAI incident, Anthropic's Claude models did not 'escape' a testing sandbox. The models were given live internet access because of a 'misunderstanding' with a third-party testing partner.

Anthropic AI Breach Explained 

After reviewing 141,006 evaluation runs where Claude could have obtained internet access, Anthropic identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, the third-party evaluation partner. The model then gained unauthorised access to the production infrastructure of three different organisations.

In each incident, Claude was tasked with a 'capture-the-flag' challenge - the model was given a fictional scenario and told that a piece of secret information (the 'flag') has been hidden on a different machine on the network, and its objective is to break in and retrieve it.

In each case, Anthropic’s prompt specified to Claude that its environment was a simulation and that it had no internet access. However, due to a misunderstanding with Irregular, internet access was made available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.

Due to the false belief that all accessible entities were intended to be part for the exercise, Claude compromised the organisations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. The model did not find or exploit any complex vulnerabilities and Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned.

However, in some cases, an older model continued its attack even after getting evidence it was running on the open internet; the most updated model stopped once it recognised it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.

Experts Warn About the Pace of AI Development 

Commenters are suggesting that major AI labs are negligent in protecting the public from their agents and that stronger government regulation is required, with guaranteed punitive damages as a deterrent.

Speaking to BBC News, cyber-security expert David Allott from Veeam Software said: "(It's) not necessarily that AI has developed a fundamentally new attack capability".

"Instead, it is that AI agents can combine capabilities, obtain credentials and system access to take actions autonomously, while adapting scope and scale at machine speed."

Mark Tilley

Mark Tilley

Mark Tilley is the Content & Community Manager for the Safety & Security Event Series, where he oversees the digital content strategy and audience engagement across a dynamic portfolio of online platforms. With a passion for connecting industry professionals and amplifying the voices shaping the future of safety and security, Mark curates and delivers thought-provoking news, interviews and insights across five publications: The Security Briefing, Health Safety Digital, Workplace Unplugged, Fire Safety Digital and Pro Integration Insider.

Author