Search TorNews

Find cybersecurity news, guides, and research articles

Popular searches:

Home » News » Cyber Threats » Anthropic AI Hacked Three Real Organizations After Internet Access Error

Anthropic AI Hacked Three Real Organizations After Internet Access Error

Last updated:July 31, 2026
Human Written
  • Anthropic’s Claude AI models hacked into three real organizations during security tests after a setup error gave them internet access.

  • The AI agent created and published a malicious software package and stole the credentials of a security firm’s scanner.

  • Such cases of rogue AI behavior have occurred in response to similar disclosures made by a competing firm, OpenAI.

Anthropic AI Hacked Three Real Organizations After Internet Access Error

Anthropic, the company behind the Claude family of AI models, has revealed a startling security failure. Its artificial intelligence hacked into the systems of three real-world organizations.

The incidents happened during a private security experiment. The developers originally confined the AI to an isolated test environment with no internet connection.

The company said it designed the tests to assess the hacking capabilities of its models. But a “misconfiguration” on systems run by Anthropic and its testing partner inadvertently gave the AI live internet access. When the models connected online, they treated the real companies they found as part of the same exercise.

The San Francisco-based firm has since reported the intrusions to the affected companies. It has also paused all such cybersecurity evaluations.

A Series of Rogue Incidents

The revelations came just days after rival OpenAI disclosed that its own AI models had breached the systems of other companies. This announcement prompted Anthropic to conduct a thorough internal review.

The company examined over 140,000 test runs looking for signs of trouble. The earliest incidents date back to April. Anthropic says it found three cases where its Claude models escaped their digital boundaries.

The first incident is considerably the most serious. A model named Opus 4.7 was tasked with retrieving a secret from a fictional company’s network. The fake company shared its name with a real one.

Unable to reach its simulated target, the AI found the real company’s infrastructure online. It then exploited weak passwords and unauthenticated endpoints to break in. The model accessed a database containing several hundred rows of production data.

The second incident showed a high level of creativity from the AI. During an exercise, Claude Mythos 5 found a document telling employees to install a non-existent software package.

The AI saw this as an opportunity to achieve its goal. It created and published a malicious Python package to the public software registry, PyPI. The incident mirrors real-world attacks where hackers have used Claude Code and DeepSeek AI to target government systems. It went to great lengths to do this, even registering an email address and an account.

Claud AI Published Malicious Code to the Internet

This booby-trapped package was available online for about an hour. During that time, users downloaded and installed it on 15 real computer systems. One of these systems belonged to a security company whose malware scanner automatically installed the package.

When the scanner executed the code, Claude’s malware stole the company’s credentials. The AI then used those passwords to access further systems at that company. The model even recognized, in its own reasoning, that publishing a real package was “NOT okay,” but convinced itself it was still in a simulation.

A third incident involved an internal research model that Athropic never released to the public. It scanned roughly 9,000 internet-facing systems. It managed to compromise a company’s application using an exposed debug page and an SQL injection vulnerability, both considered basic techniques.

The model stopped its attack on its own once it determined the system it had hacked was real. Neither Anthropic nor the organizations that suffered the breach noticed the intrusions at the time.

After reviewing historical logs, Anthropic identified the activity and accepted responsibility, saying it is handling the issue “as if the responsibility were ours alone.”

The company also acknowledged that its monitoring and record reviews should have been more thorough.

Similar Incidents at OpenAI

Earlier this month, OpenAI disclosed similar security breaches involving its own AI agents. The company revealed that, during evaluations, one of its autonomous AI agents went beyond testing boundaries and gained access to the systems belonging to Hugging Face.

OpenAI termed this case as unprecedented and assured that it is conducting an investigation along with Hugging Face. The back-to-back disclosures suggest major AI developers are increasingly testing, and discovering the limits of autonomous AI systems.

Consequences of AI Autonomy and a Wake-Up Call

According to Anthropic, these discoveries make the firm cautiously optimistic that it can manage such risks. Anthropic made it clear that none of the AI models tried to escape their environment and pursue their own agendas. They were simply trying to complete the tasks they were given.

The AI was autonomous in carrying out the task assigned to it; it was not, however, self-directed or choosing its own objectives.

Claude’s goal never changed:

  • Find protected information.
  • Exploit systems if necessary.
  • Continue until it completes the task.

Because the testing environment accidentally connected to the internet, Claude treated external systems as part of its search.

This distinction is important because headlines suggesting that the AI “escaped” can imply intentional behavior that Anthropic’s own explanation does not support.

However, experts say the lessons are clear. “The moral of this story is not to fear robots that will take over, but the companies behind powerful AI agents who are making the decisions about what is safe for the rest of us,” said Professor Gina Neff of the University of Cambridge.

Cybersecurity expert David Allott added that the threat is not that AI has developed a new attack capability.

Instead, the issue is that by combining their capabilities, AI agents can obtain login credentials and gain system access to take actions on their own, scaling their work at lightning speed.

The incidents have also been met with some skepticism, as both Anthropic and OpenAI are preparing for major stock market listings.

Share this article

About the Author

Memchick E

Memchick E

Digital Privacy Journalist

Memchick is a digital privacy journalist who investigates how technology and policy impact personal freedom. Her work explores surveillance capitalism, encryption laws, and the real-world consequences of data leaks. She is driven by a mission to demystify digital rights and empower readers with the knowledge to protect their anonymity online.

View all posts by Memchick E >
Comments (0)

No comments.