Anthropic CEO Calls for Slowing Down AI Development Process
Anthropic executive Dario Amodei shared a three-stage strategic plan aimed at increasing audits and developing safety measures in the artificial intelligence sector.
Anthropic CEO Dario Amodei stated that the pace of development of artificial intelligence technologies should be slowed down and announced a three-stage plan that includes granting external auditors access to their models.
Plan to Slow Down AI Development
Anthropic CEO Dario Amodei proposed a comprehensive three-stage plan to slow down the pace of artificial intelligence work and allow necessary safety measures to be put in place.
This approach aims to buy time for companies to establish safety measures and for regulators to evaluate the models.
External Auditors' Access to Modeling
As part of the first step of the plan, Anthropic is unilaterally implementing a policy to grant third-party evaluators, such as METR, broad access to its models.
This step is of critical importance in terms of verifying compliance with safety practices and commitments.
Sectoral and Democratic Cooperation
In the second stage, the sector is planned to come together with government agencies to establish common safety standards and set limits on uncontrolled progress.
It is emphasized that because legal processes take time, companies in democratic countries should work together now to establish standards.
Global Standards and Authoritarian Regimes
The third and most challenging stage involves convincing authoritarian regimes like China and Russia to slow down development as well and adopting global standards.
During this process, it is stated that access to high-powered chips should be restricted and imitation attempts prevented in order for the US and its allies to maintain their technological superiority.
Concerns Over Recursive Self-Improvement
At the root of Amodei's concerns lie recursive self-improvement processes, where artificial intelligence systems train subsequent generations themselves.
It is stated that if this situation is not kept under control, the capacity to understand and manage the systems may remain inadequate.
Security Incidents and Risks
It was recalled that during the summer incidents involving OpenAI and Hugging Face, agent swarms attempted to launch unintended cyberattacks and tried to hack the evaluator system.
Additionally, Anthropic's Claude model had also come to the agenda with similar cybersecurity incidents in the past period.