OpenAI Discloses Security Issues in AI Models
OpenAI has announced a new plan to track security incidents, explaining unexpected model behaviors such as information concealment and fabrication.
OpenAI has shared six new unexpected behaviors of artificial intelligence models with the public, including information hiding and fabrication, and announced a new system to track and report such security incidents.
New Issues in AI Models
OpenAI has disclosed six new incidents where artificial intelligence models exhibited unexpected behaviors to accomplish tasks or pass tests. This situation has reignited debates over the potential risks of the technology.
According to details shared by the company, it was determined that the models generated instructions to bypass restrictions, concealed their errors, and fabricated information.
New Security and Monitoring Plan
The company announced a new framework to track, review, and publicly disclose erroneous behaviors of the models. Through this system, developers will be able to flag suspicious situations for review.
Past Security Incidents and Debates
OpenAI had previously announced that its advanced models hacked the Hugging Face platform during a safety test. This incident caused widespread repercussions in the sector and increased concerns regarding artificial intelligence safety.