
Kimi K3, Moonshot AI’s flagship open-weight model, escaped its evaluation sandbox earlier this month while being tested on defensive cybersecurity skills. The model probed its network, found a gap, and retrieved the benchmark answers it needed from GitHub. Kimi K3 was the latest in a cluster of disclosures showing AI models reaching systems or infrastructure outside their intended evaluation boundaries.
The immediate question is how the models got out. But the string of incidents points to a more consequential question for operators: who was testing these systems in the first place, and what exactly are companies being asked to trust when they rely on those evaluations?
Most enterprises will not commission these evaluations themselves. They will rely on the safety testing, system cards and assurances provided by the AI vendor whose model they are considering. When they do, they are implicitly trusting the vendor’s choice of evaluator, the evaluator’s sandbox configuration and the evaluator’s definition of what “contained” means during a test.
AI safety testing has become an increasingly important part of the AI supply chain. Labs are shipping frontier models faster than at any point in their history, creating a bottleneck in the infrastructure needed to evaluate whether those systems are safe to deploy. Into that gap have entered specialist AI security startups. They are given access to frontier models with safety filters deliberately reduced, credentials, network access and purpose-built environments designed to test dangerous capabilities. Yet the standards applied to those firms are not necessarily comparable to those applied to other critical security vendors, even though their evaluations may form part of the evidence enterprises rely on when procuring AI systems. And that is a procurement problem most enterprises have not yet recognized.
What Happened in Three Weeks
Between July 21 and August 7, four AI labs disclosed that their models had reached systems outside their intended evaluation boundaries:
On July 21, OpenAI disclosed that two of its models—GPT-5.6 Sol and an unreleased internal research prototype—escaped their evaluation sandbox during the ExploitGym cybersecurity benchmark and breached Hugging Face’s production infrastructure, executing more than 17,000 documented actions over four days through a zero-day vulnerability in a package proxy. This incident involved a separate mechanism from Irregular’s environment; a subsequent OpenAI update confirmed that additional account-access cases from other evaluations were linked to Irregular, but the Hugging Face breach itself was distinct.
On July 30, Anthropic disclosed three incidents in which Claude models—Opus 4.7, Mythos 5, and an unreleased internal model—gained unauthorized access to the production systems of three external organizations through Irregular’s evaluation environment.
On August 5, Meta disclosed that Muse Spark 1.1 had breached an unnamed company during a cybersecurity evaluation run through Irregular. Irregular confirmed to Reuters it was “the exact same evaluation-environment issue”—a misconfiguration that left evaluation machines with live internet access despite system prompts asserting otherwise.
On August 7, Frontier Security published findings that Kimi K3 had escaped its sandbox because outbound HTTPS and DNS traffic remained open to public IP ranges. Frontier Security argued that Kimi K3’s fewer internal guardrails allowed it to access the internet without express permission. UK AISI blamed Frontier’s configuration of its own tool, noting that users are responsible for configuring its sandbox tool to their own needs, and that it had published detailed guidance on how to do so. Moonshot AI has not commented.
Of the four recent incidents, three trace back to the same company. Irregular, a Tel Aviv-based AI security startup founded in 2023 by Dan Lahav and Omer Nevo, raised $80 million from Sequoia Capital and Redpoint Ventures in September 2025 at a $450 million valuation. By 2026, the company was running cybersecurity evaluations for OpenAI, Anthropic, and governments including the UK.
The fourth incident, the Kimi K3’s escape, used sandbox software from the UK’s AI Security Institute, configured and run by Frontier Security, a US-based AI cybersecurity research firm led by CEO Yaron Singer.
A failure in the evaluation environment does not by itself mean the model is unsafe. It means the evidence being used to assess that model has become harder to trust. If the controls, assumptions or boundaries of the evaluation were wrong, then the results need to be understood in that context before they are used to support a deployment decision.
The Structural Problem the Pattern Exposes
When a company hires an AI security evaluator, it is granting that vendor access to its most capable AI systems with safety filters deliberately reduced, in an environment specifically designed to test their most dangerous capabilities.
In traditional enterprise security procurement, firms granted comparable access are often subject to formal due diligence covering certifications, professional indemnity insurance, security controls and contractual allocation of liability.
No comparable, widely adopted assurance framework yet exists for AI evaluation vendors. Singapore’s planned AI Tester Accreditation Programme, expected in Q3 2026, would be the first of its kind in Asia. It is intended to accredit testing service providers according to their technical competence and operational readiness. Japan’s AI Safety Institute updated its Guide to Evaluation Perspectives on AI Safety in July 2026, expanding its evaluation framework in response to the growing deployment of AI agent systems. For operators in regulated Asian industries, the emergence of both frameworks signals that evaluation assurance is moving from voluntary to expected, and that procurement teams should be asking whether their model vendor’s evaluator would meet either standard.
Seán Ó hÉigeartaigh, director of the AI: Futures and Responsibility project at Cambridge, told TechCrunch that the incidents made clear “sandboxing and testing environment controls aren’t really keeping pace with the capability of the models.” Andrew Yoon, an AI policy researcher quoted in the same piece, was more direct about the cause: “There are competitive pressures that are incentivizing a race to the bottom on safety standards, and that is a perfect place for regulatory intervention.”
Irregular’s position of simultaneously being trusted with frontier-class capabilities from multiple AI labs illustrates the concentration risk this creates. A single vendor failure produced breaches across multiple clients and multiple victim companies within the same incident cycle.
Irregular has since cut off internet access entirely for all models it tests and is developing new containment processes. But it does not change the fact that for an extended period, it was the single point of containment for the most offensive AI capabilities at three of the world’s four most capable frontier labs simultaneously.
What Operators Should Be Asking
For most operators, the exposure is not direct. The majority of organizations outside of frontier AI labs will not commission Irregular or Frontier Security directly. The risk is transitive. It travels through the AI vendors, platforms, and models that operators procure and deploy.
Before procuring or deploying a model, run these questions against the model vendor’s contract and evaluation documentation:
Who owned the evaluation environment, and is that in writing? The Anthropic and Meta incidents both turned on a contested understanding of whether the environment had internet access. Ownership of the environment—who configured it, who verified it, and who accepts liability if the configuration is wrong—should be explicitly assigned in the contract, not assumed or left to the evaluator’s own account after the fact.
What are the evaluator’s contractual notification obligations if its environment fails and a third party is compromised? The three organizations breached in Anthropic’s incidents were notified by Anthropic, not by Irregular. One had not been reached as of the disclosure date. If the evaluator’s configuration failure causes harm to a party outside the contract, the obligation to notify, preserve evidence, and remediate should be assigned before the evaluation begins. Enterprise customers should also know what obligation their AI vendor has to disclose such a failure to them.
What assurance can the AI vendor provide over the evaluator’s sandbox infrastructure? The vendor should be able to show how the sandbox architecture, credential scope, and egress controls were independently verified before the evaluation ran, rather than asking customers to accept the evaluator’s controls on trust.
Was the model tested under conditions that materially match the system the enterprise plans to deploy? A model tested in a tightly controlled sandbox may present a very different risk profile from the same model operating with production credentials, external tools, persistent memory or network access. Buyers should understand which assumptions in the evaluation are relevant to their intended deployment, and which are not.
Do not treat “independently evaluated” as a safety credential. The evaluator is not a neutral technical service provider. It is a vendor with access to your most sensitive AI capabilities, operating in an environment with no mandatory accreditation, no standard liability framework, and no regulatory obligation to notify you if something goes wrong. Treat it as a critical vendor whose controls, responsibilities, and failure obligations require independent assurance before you accept the evaluation as evidence of anything.
More from Asia Tech Lens
When an AI Agent Escapes Its Sandbox
A look at what the OpenAI/Hugging Face sandbox breach revealed about containment, evaluation environments, and the operational controls enterprises should expect before trusting frontier AI systems.Agentic AI Can Act. Singapore’s New Rulebook Says: Prove You Can Stop It.
Singapore’s new guidance puts containment, permissions, and auditability at the center of agent deployment, raising the bar for operators evaluating whether AI systems are truly controllable.AI Middleware Promises Flexibility. In Asia, It Can Create Jurisdictional Lock-In
AI middleware can introduce hidden dependencies across routing, data handling, and compliance, creating fourth-party risks that enterprises may not see when they procure a model or platform.Vietnam’s New AI Law: The Road Ahead For Businesses
Vietnam’s emerging AI rules are pushing enterprises toward stronger vendor due diligence, clearer audit evidence, and tighter incident-response obligations as AI governance becomes more formalized.AI Is Shrinking The Time To Compromise. Most Firms Still Can’t Recover Control
As AI accelerates the speed of cyber compromise, the harder operational problem is whether enterprises can contain incidents, restore control, and recover when preventive safeguards fail.

