Gemini’s breach of real companies exposes an AI guardrail problem
Cybersecurity Classified by Officially
Google says one of its Gemini models accessed systems belonging to three real companies during a cybersecurity evaluation in May.
The model reportedly guessed credentials in one case, while finding exposed credentials in public repositories in two others. Google says Gemini stopped once it recognized that it had reached real infrastructure and that the affected organizations were notified.
Gemini was participating in an evaluation run by Irregular, a third-party AI cybersecurity testing firm. Similar incidents involving models from Anthropic, OpenAI, and Meta have also been linked to the same underlying problems with evaluation environments that allowed the models to reach the public internet.
But the news arrives at a particularly interesting moment.
For months, the AI industry has been steadily escalating its demonstrations of agentic capability: Models can browse, use tools, write code, pursue multi-step goals, and sometimes find ways around obstacles their developers did not anticipate. The market rewards eye-catching evidence of autonomy. “It completed the task” is impressive. “It did something it was not supposed to do” can be even more memorable.
There is another way to look at the incident: not as evidence of an AI suddenly developing criminal intent, but as a small, concrete example of the “AI alignment” problem.
Alignment is the deceptively difficult task of making an AI system’s behavior match what people actually intended, rather than merely the narrow objective they managed to express. In this case, the objective was to locate hidden information within a simulated target and complete the evaluation. But any human operator would likely have regarded one condition as non-negotiable: Do not attempt to access real companies.
This is an extract. The publication continues at the source.
Read the original at the source: https://www.malwarebytes.com/blog/ai/2026/09/geminis-breach-of-real-companies-exposes-an-ai-guardrail-problem
Officially imported this from Malwarebytes’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.
Provenance
- Organization
- Malwarebytes — imported from official source
- Official source
- https://www.malwarebytes.com/blog/feed/index.xml RSS
- Imported
- September 21, 2026 14:30
- Versions
- 1 recorded
- Identity
https://www.malwarebytes.com/blog/ai/2026/09/geminis-breach-of-real-companies-exposes-a...