Google’s Gemini AI model accessed the systems of three real companies during a cybersecurity exercise after it was unintentionally given internet access, according to a Wall Street Journal report. The incidents add to growing concerns about how increasingly capable AI systems can behave when given access to online resources and security tools.

Irregular, a company that carries out security testing and red-team exercises for AI developers, conducted the tests. Google said the incidents occurred as part of a controlled security exercise and that Gemini stopped after recognising it had reached real companies’ systems.

Gemini gained access to real systems

According to The Wall Street Journal, the Gemini model was asked to identify weaknesses in a fictional company within a contained testing environment. However, the model unintentionally gained internet access during the exercise, allowing it to interact with systems outside the intended environment.

The model reportedly gained access to three separate companies in different ways. In one incident, it guessed passwords until it successfully entered a protected system. In the other two cases, it discovered credentials in a publicly accessible repository and used them to access protected systems.

“In one of the cases, the model guessed passwords until it gained access to a protected system. In the other two cases, the model found credentials in a public repository that allowed it to access protected systems then. In each case, the model ended the intrusion after determining it had accessed a real company’s systems,” the WSJ report said, citing Google.

The incidents appear to have resulted from the model interpreting the security task more broadly than intended. Although the exercise was designed around a fictional organisation, Gemini encountered real-world infrastructure associated with companies sharing the same name. Google said the model stopped once it recognised that it was dealing with real systems rather than continuing its activity.

The identities of the three companies have not been disclosed. Google said it informed the affected organisations.

Google says Gemini behaved appropriately

Google has described the incidents as part of a bug bounty-style cybersecurity exercise rather than a security breach requiring public disclosure. The company said the model stopped its actions once it determined that it had reached genuine company infrastructure.

“In this case, the model acted appropriately,” a Google spokesperson said.

Irregular has previously conducted red-team security work involving AI models from several major technology companies, including OpenAI, Anthropic and Meta. Such exercises are designed to identify potential weaknesses before similar behaviour can occur outside controlled testing environments.

The Gemini incidents are notable because the model was not supposed to have internet access. Irregular said that “internet access was unintentionally made available”, creating an unexpected path for the model to interact with external systems.

The distinction between deliberate security testing and unintended real-world access is important. The companies were not reportedly targeted in a conventional criminal attack, and the model stopped after recognising the nature of the systems it had reached. Nevertheless, the incidents show how a small change in an AI system’s environment can have significant consequences.

AI security tests highlight growing risks

The Gemini incidents come as researchers and AI companies increasingly test advanced models’ ability to perform complex cybersecurity tasks. Recent tests involving models from other AI developers have shown that systems can identify vulnerabilities, search for credentials and carry out multi-step actions with limited human intervention.

Anthropic’s Claude has also been involved in security research examining its ability to penetrate computer systems. In one recently reported case, researchers used Claude during an authorised test involving OpenAI. Other incidents involving AI agents have raised questions about how models could behave when given broader access to networks, code repositories and other online resources.

These developments have placed greater attention on the concept of AI alignment, which broadly concerns whether an AI system behaves according to its intended objectives and constraints. A model carrying out an unexpected action does not necessarily mean that it is misaligned, however, and Google has rejected the suggestion that the Gemini incidents represented a case of model misalignment.

Instead, the company has characterised the behaviour as appropriate within the circumstances because Gemini stopped after discovering that it had reached real systems. The episode nevertheless illustrates the challenges involved in designing safeguards for AI systems that can independently search for information and take actions across connected environments.

As AI agents become more capable, developers are increasingly testing them in simulated environments before allowing them to interact with real infrastructure. The Gemini incidents show why those boundaries can be hard to maintain when an AI system has internet access or encounters unexpected information during a task. The three affected companies have been notified, but the organisations’ identities and further technical details remain undisclosed.

Share