Debate on Independent Safety Evaluators in AI Companies
While the proposal by Anthropic and OpenAI leaders to embed third-party safety evaluators into AI companies has been welcomed by the sector, uncertainties regarding independence and data access draw attention.
Anthropic and OpenAI executives have announced plans to integrate independent third-party evaluators into their operations to oversee the safety of pioneering AI systems. While experts find this step positive, they emphasize that legal guarantees are essential for complete independence.
Details of the Safety Evaluator Proposal
Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman proposed embedding independent safety evaluators within frontier AI companies. This step aims to assess model alignment, report incidents, and share findings.
The Approach of Independent Research Groups
Independent research groups such as METR, Redwood Research, Apollo Research, Far.AI, Palisades Research, and Safer AI welcomed the proposal. However, these groups stated that past constraints regarding data access and time must be overcome for success.
Access Demands and Intermediate Milestones
Experts draw attention to the importance of gaining access not only to finished models but also to intermediate training stages. Far.AI CEO Adam Gleave stated that this would make it clearer when concerning behaviors emerge.
Limitations of Past Evaluations
In past tests, strict non-disclosure agreements and short timeframes applied by companies negatively impacted the evaluations. It is known that organizations such as Apollo Research and METR were granted very limited timeframes in the past.
Legal Regulations and Future Expectations
Stating that voluntary measures depend on the goodwill of companies, experts demand transparent public frameworks and mandatory regulations. Laws in California and the EU AI Act offer new frameworks in these processes.