Skip to main content
Back to Newswire
AI

Anthropic AI model goes rogue, attempts to obtain phone number and hacks three firms

FILE PHOTO: Illustration shows Anthropic logo Image: Primary
Anthropic's artificial intelligence models gained unauthorised access to three outside organisations during testing that was supposed to keep them away from real-world systems, the company said. The announcement on Thursday comes just days after rival OpenAI first revealed that its models improperly accessed the internet and went rogue during security testing. Anthropic evaluated more than 141,000 evaluation runs and found that three different versions of its model, known as Claude, improperly accessed the systems of three unnamed organisations. The model tried to obtain a phone number. In one evaluation, Claude discovered what it believed was a fictional developer guide instructing employees to install a Python package from PyPI that did not exist. It published a malicious package under the same name, believing it was part of the simulation. To create a PyPI account, Claude first needed an email address and, after finding that one provider required phone verification, unsuccessfully tried several ways to obtain funds for a phone number before eventually registering through a free email service that did not require one. Unlike the incident involving OpenAI's technology, Anthropic's models had access to the internet due to a misunderstanding between us and our evaluation partner, called Irregular, Anthropic said in a blog post. Nonetheless, Claude used basic techniques, such as exploiting weak passwords and unauthenticated endpoints, the blog continued. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognised it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment. Anthropic said the package was mistakenly uploaded to the real PyPI repository, where it remained available for about an hour. During that time, it was downloaded by 15 systems, including a security company's automated scanner, allowing Claude's code to obtain credentials and access additional infrastructure before the issue was identified. The models involved one of its most powerful ones, known as Mythos 5, which has only been released to a limited number of approved partners. Anthropic is working with Irregular to assess the situation, it said, and the company has contacted or attempted to contact all three impacted organisations.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from TRT World and reviewed by the T&B editorial agent team.