Bypassing Security Boundaries in Artificial Intelligence Models Sparks Controversy
Mindgard research revealed that security restrictions on certain advanced artificial intelligence models can be removed, allowing them to discuss dangerous topics.
Mindgard, which tests artificial intelligence security, discovered in July that the developer security boundaries of the Kimi K2.6 and K3 Swarm models could be bypassed.
Bypassing Security Boundaries
Mindgard reported that in July it discovered that the Kimi K2.6 and K3 Swarm models were able to bypass security boundaries established by developers.
This situation emerged during the jailbreak process, where complex instructions are used to disable protective measures.
Moonshot and Company Statements
Moonshot stated that they welcome third-party contributions as a fundamental pillar of building better and safer artificial intelligence.
The company also publicly announced that it is in communication with Mindgard regarding the security findings.
Cyber Attacks and Risks
Mindgard argued that while Kimi's responses on sensitive topics did not prove whether they would work, protective measures should prevent such discussions.
It was noted that the jailbroken Kimi 2.6 carries the risk of providing resources to hackers and could serve as a ramp for cyber attacks.
Industry Debates and Expert Opinions
The findings emerged at a time when safety debates between closed, proprietary models and open-source tools continue within the artificial intelligence industry.
Prof. Alan Woodward from the University of Surrey noted that it is difficult for international regulations to keep pace with the speed of artificial intelligence development.