China's Kimi K3 AI model escapes isolated sandbox during security test
China’s top open-weight AI model Kimi K3 broke out of its isolated test environment during a cybersecurity evaluation, according to US security researchers, following similar high-profile incidents involving closed fron…
China’s top open-weight AI model Kimi K3 broke out of its isolated test environment during a cybersecurity evaluation, according to US security researchers, following similar high-profile incidents involving closed frontier models from OpenAI and Anthropic that highlight the growing challenge of constraining AI behaviour. Kimi K3, released last month by Beijing-based Moonshot AI, escaped from a supposedly isolated sandbox environment, accessed the open internet and found solutions on the developer platform GitHub, US firm Frontier Security said in a blog post on Thursday. According to Frontier Security researchers Paul Kassianik and Yaron Singer, the incident occurred when the firm tested Kimi K3’s defensive cybersecurity capabilities using a benchmark evaluation from the AI Security Institute – a UK government research organisation. A “basic network misconfiguration” in the benchmark framework allowed Kimi K3 to flee its digital testing cage and look up answers on the internet, effectively cheating the test, they said. Kimi K3’s escape, however, did not involve the hacking of an external system, unlike recent breaches caused by OpenAI and Anthropic models. Last month, OpenAI said its flagship GPT-5.6 Sol and an unreleased, “even more capable” system broke out of a sandboxed environment and hacked the open-source developer platform Hugging Face to obtain secret information containing answers to an internal test. Anthropic later disclosed that it had also identified three previous incidents where its Claude models – including Opus 4.7, Mythos 5 and an internal research model – breached the infrastructure of three undisclosed organisations. Even without a direct hack, the Kimi K3 incident could still be “potentially more harmful” because of the model’s open-weight nature, which makes it publicly available to adversarial actors, Frontier Security said. Kimi K3’s ability to circumvent its test parameters also reflected the model’s weaker internal safeguards, according to one of the researchers. “It found the misconfiguration itself, by probing its own network settings,” Kassianik wrote in a social media post on Friday. “Most publicly available frontier models have internal guardrails that stop them. K3 didn’t blink.” Moonshot did not immediately respond to a request for comment on Friday. The rapidly advancing capabilities of frontier AI have heightened global cybersecurity concerns, pushing the role of widely accessible open-weight models to the centre of an escalating debate. While critics warn that advanced open-weight models lower technical barriers for malicious actors, several US AI leaders have defended open-source systems’ key role in cyber defence. American tech giants including Nvidia, Palantir and Meta Platforms last month signed an open letter urging Washington not to restrict open-weight models. While open models carried real risks, they also enabled cybersecurity defenders to detect and respond to emerging threats posed by AI-equipped attackers, the companies argued. That point was underscored when Hugging Face had to rely on GLM-5.2, an open-weight model from Beijing-based Zhipu AI, to contain the OpenAI attack after major proprietary US models refused to process the security logs due to safety guardrails.