Anthropic has announced the release of their Responsible Scaling Policy (RSP), designed to manage the risks associated with developing advanced AI systems. As AI models gain more capabilities, they promise significant economic and social benefits but also pose severe risks. The RSP mainly addresses catastrophic risks, where AI might cause extensive destruction either through misuse by malicious actors or unintended autonomous actions.
The policy introduces a framework known as AI Safety Levels (ASL), conceptually similar to the US government’s biosafety level standards for handling hazardous biological materials. Each ASL level demands specific safety, security, and operational protocols depending on the AI system’s potential catastrophic risks.
The ASL framework is summarized as follows:
1. ASL-1 includes systems with no significant catastrophic risk, such as a 2018 language model or an AI that plays chess.
2. ASL-2 covers systems showing early signs of dangerous capabilities, like instructing on bioweapon creation, but the information isn’t reliably actionable.
3. ASL-3 concerns systems that markedly elevate catastrophic misuse risk in comparison to non-AI baselines or exhibit low-level autonomous behaviours.
4. ASL-4 and higher levels are not yet defined but will involve more stringent safety criteria as AI capabilities advance.
Currently, ASL-2 represents Anthropic’s existing safety and security measures, paralleling commitments made to the White House. ASL-3 introduces tighter standards requiring extensive research and engineering, such as robust security protocols and red-team adversarial testing. ASL-4 measures remain undefined, aiming to address advanced unsolved research challenges.
The ASL framework balances minimizing catastrophic risks while fostering beneficial applications and progress. By incentivizing solutions to safety issues, it encourages advancements in AI while temporarily pausing further training if safety protocols lag behind. This approach could motivate a “race to the top” among AI developers to enhance safety standards.
Anthropic clarifies that the RSP will not affect current usage or availability of Claude, their AI product. Instead, this policy functions like pre-market testing in the automotive or aviation industries, ensuring rigorous safety standards before new AI models are released, ultimately benefiting users.
The board has approved the RSP, and any changes require further board consultation with the Long Term Benefit Trust. Anthropic acknowledges that rapid AI field developments necessitate continual iteration and adjustment of the policy.
The full details of the Responsible Scaling Policy can be found [here](https://www.anthropic.com/responsible-scaling-policy). Anthropic hopes it will inspire policymakers, nonprofit organizations, and other companies dealing with similar AI deployment challenges.
Acknowledgments were given to ARC Evals for their critical insights and expertise, particularly in AI risk assessment and developing evaluation procedures.
You can read the original article here: [https://www.anthropic.com/news/anthropics-responsible-scaling-policy](https://www.anthropic.com/news/anthropics-responsible-scaling-policy).