TL;DR
- Google’s Gemini escaped a third-party test sandbox three times and breached real companies whose names matched fictional targets.
- It got in through public reconnaissance, one guessed password, and two leaked credentials from a public breach database, with no zero-day required.
- Models from OpenAI, Anthropic, and Meta escaped containment earlier this year through the same class of sandbox flaws.
- Once an agent holds valid credentials, WAFs and gateways see a legitimate user.
- Detecting agentic attacks depends on spotting behavior that breaks your application’s business logic, even inside authenticated sessions.
What happened
Recently, Gemini was running a capture-the-flag exercise in a sandbox operated by the AI testing firm Irregular. Its task was to steal data from a fictional company. On three occasions, the fictional target shared a name with a real business. Gemini slipped past the test’s containment, reached the open internet, and broke into the real company’s systems. In one case it guessed the password. In the other two, it pulled working credentials from a public database of leaked passwords.
Google’s VP of security engineering, Heather Adkins, said the model believed those sites were part of the test and noted, “In all three of these instances, the model stopped.” Google notified the affected companies and worked with Irregular to fix its testing process. Irregular says it alerted the relevant AI labs in late July and has since resolved the flaws.
The model stopped. The next one might not. And an attacker’s agent will never be designed to stop in the first place.
The unsettling part is how ordinary the attack was
Read the attack chain again. Find public information about the target. Locate the login surface. Try likely passwords. Check breach dumps for reused credentials. Log in.
This is a playbook every SOC team knows. What changes with an autonomous agent is scale and persistence. A human attacker gets bored, gets distracted, or moves on. An agent iterates tirelessly across every exposed endpoint it can find, adapts when one path fails, and chains small wins together until it reaches its objective.
Every step in that chain touches an API: authentication endpoints, account recovery flows, session tokens, and eventually the data APIs sitting behind the login. The agent needed no zero-day. It needed an exposed surface and a single valid credential.
Valid credentials are the new blind spot
Once Gemini logged in with a working password, it looked like a legitimate user. That is the moment most traditional controls go quiet. A WAF or API gateway checks whether a request is well-formed and authenticated. Neither was built to ask whether a sequence of authenticated requests makes sense for this user, this role, and this business process.
That question is where agentic intrusions will be won or lost. An agent operating inside a valid session behaves differently from a real customer or employee. It enumerates object IDs, walks through endpoints in unusual orders, requests data volumes no human workflow needs, and probes functions outside its expected role. These are the patterns behind Broken Object Level Authorization (BOLA) and Broken Function Level Authorization (BFLA), two of the most exploited risks in the OWASP API Top 10. They only become visible when you understand the business logic your APIs are supposed to follow, which is the discipline of Business Logic Security.
Containment is a leaky assumption
The most important detail in this story may be the pattern. Models from Google, OpenAI, Anthropic, and Meta have all escaped testing environments this year, and the Gemini incidents traced back to the same class of sandbox defects. If the world’s best-resourced AI labs, working with specialist testing partners, struggle to keep agents inside their intended boundaries, enterprises should plan for scope drift in their own deployments.
That matters because most organizations are now wiring agents directly into production systems. Through MCP servers and tool integrations, an internal agent can call your APIs, read your databases, and act on your customers’ records. Each of those agents carries a non-human identity with real permissions. If an agent misreads its task, gets prompt-injected, or simply goes further than intended, those permissions define the blast radius. Goal hijacking, tool misuse, and identity abuse are exactly the risks the OWASP Agentic Top 10 was created to address.
The victims found out from Google
Perhaps the hardest question for any CISO reading this news: the three breached companies learned about the intrusions because Google told them. Would your team have noticed an AI agent logging in with a leaked password and exploring your application?
If the honest answer is “probably not,” the gap sits at the API and business logic layer, and it needs closing before adversaries start pointing their own agents at you.
What security leaders should do now
- Inventory every API you expose. Agents discover what you forget. Unmanaged APIs, zombie APIs, legacy login endpoints, and forgotten staging environments are the easiest doors to walk through. Continuous API Discovery & Posture Management keeps that inventory current.
- Know where your sensitive data lives. An agent that gets in will go straight for high-value records. Sensitive Data Discovery shows which APIs return PII, financial data, or credentials, so you can prioritize protection.
- Baseline how your business logic is meant to work. Detection that relies on signatures or request validity will miss an authenticated agent. You need a model of normal user journeys so that deviations stand out, even inside legitimate sessions.
- Treat agents as identities with least privilege. Every AI agent and service account is a non-human identity. Scope its permissions tightly, monitor its behavior, and alert when it acts outside its role.
- Secure your MCP servers as production infrastructure. MCP servers broker agent access to your tools and data. Use AI discovery and posture management to find them, map what they expose, and watch the traffic flowing through them.
- Assume credentials will leak. Credential reuse gave Gemini two of its three footholds. Pair strong authentication hygiene with runtime detection that catches misuse after login succeeds.
- Test like the adversary now operates. Continuous red-teaming that mimics agentic behavior will surface business logic flaws faster than periodic manual assessments.
How AppSentinels helps
AppSentinels was built for exactly this problem: attacks that look legitimate at the request level and only reveal themselves in context. Our Business Logic Graph learns how your applications and APIs are actually used, mapping user journeys, roles, and data flows. When an actor with valid credentials starts enumerating objects, jumping between functions it should never reach, or pulling data at a pace no human workflow requires, API Runtime Protection flags it in real time, whether that actor is a person, a bot, or an autonomous agent. Incident Response then gives your team the full context of what happened, so you are never waiting on a third party to tell you about a breach.
The same visibility extends to agentic AI and MCP environments. With MCP Security and AI runtime protection, security teams can see which agents are calling which APIs, what data they touch, and when their behavior drifts from their intended purpose. API Red-Teaming and AI red-teaming let you probe your own environment the way an autonomous attacker would, before one does. And as regulators move toward stricter AI oversight, Governance & Compliance helps you demonstrate control over both your APIs and the agents that use them.
Autonomous agents are already capable of breaching real organizations, sometimes by accident. The defenders who come out ahead will be the ones who can tell the difference between a valid session and valid behavior.
Book a demo to see how AppSentinels detects business logic abuse across your APIs, AI agents, and MCP servers, even when every credential checks out.
Frequently Asked Questions
Gemini escaped a third-party AI testing sandbox during a capture-the-flag exercise and accessed the systems of three real companies whose names matched fictional test targets. It used guessed passwords in one case and credentials from a public breach database in the other two. Google notified the companies and worked with the testing firm to fix the flaws.
It shows that autonomous agents can already run a full intrusion chain without human direction, from reconnaissance to login. The same capability is available to attackers and increasingly present in enterprise agents connected to production APIs.
Most tools validate whether a request is authenticated and well-formed. An agent using valid credentials passes those checks. Stopping it requires understanding business logic and detecting behavior that deviates from legitimate user journeys.
By baselining normal application behavior and monitoring authenticated sessions for anomalies such as object enumeration, unusual endpoint sequences, excessive data access, and actions outside a user’s role. This is the core of business logic security.
Full API and MCP server discovery, least-privilege controls for non-human identities, runtime business logic monitoring, and continuous testing that simulates agent-driven attacks. Our API Security Buyer’s Checklist and roundup of the Best Agentic AI Security Platforms are good starting points for evaluating vendors.