Security Researchers Infiltrated OpenAI Systems Using Anthropic Model
Independent security experts from the Hacktron AI startup combined critical vulnerabilities to gain access to OpenAI infrastructure and earned a bounty.
A three-person security team from the Hacktron AI startup successfully bypassed OpenAI's defense mechanisms by using Anthropic's artificial intelligence model. By chaining critical security vulnerabilities, the researchers gained access to employee accounts and won a $6,500 bug bounty.
AI-Powered Security Breach
Independent security researchers bypassed OpenAI's defense systems and gained access to the company's infrastructure using Anthropic's AI model. This development highlights the boundaries of a new era in AI security.
A three-person security team from the Hacktron AI startup carried out this attack as part of OpenAI's bug bounty program. After the researchers reported their findings, the company awarded the team a $6,500 bounty.
Chaining Two Critical Vulnerabilities
The researchers successfully combined two critical vulnerabilities to access the ChatGPT accounts of multiple OpenAI employees. This method enabled attackers to directly log into corporate software.
OpenAI officials announced that these security flaws uncovered by the Hacktron team have been completely resolved. The incident coincided with a period when leading AI companies face intense security pressures.
The Initial Entry Point in Discourse Software
On July 25, researchers targeted the system by exploiting a flaw in Discourse, the third-party software powering the infrastructure of OpenAI's community forum. The entry point was established via the upload of an ordinary image in HEIF or HEIC format.
A memory corruption error that occurred during the processing of images via ImageMagick and libheif paved the way for attackers to inject their own instructions into the system.
The Role of the Anthropic Model and Rapid Fix
The researchers noted that the Opus model they initially used failed to generate a working exploit code. However, this changed with the release of a new version by Anthropic, and the attack was successfully completed.
The team that infiltrated the Discourse server found a second vulnerability enabling the takeover of employees' ChatGPT and Codex accounts. After the researchers reported the situation, Discourse and OpenAI released the necessary patch on July 27.