Anthropic has introduced a Responsible Scaling Policy (RSP) to manage the risks posed by developing increasingly capable AI systems. As AI models evolve, they have the potential to create significant economic and social value but can also present severe risks. The RSP focuses on catastrophic risks, where an AI model could cause large-scale devastation. Such threats might arise from the deliberate misuse of AI models or from AI systems acting autonomously in unintended ways.
The policy introduces the AI Safety Levels (ASL) framework, which aims to address catastrophic risks. This system, inspired by U.S. government biosafety standards, imposes progressively stricter safety, security, and operational standards based on an AI model’s potential for catastrophic risk. ASL-1 includes systems posing no significant risk, such as a 2018 language learning model (LLM). ASL-2 encompasses systems with early signs of dangerous capabilities, although not yet highly reliable. Both present AI models, including Claude, fall under this category. ASL-3 escalates the standards for systems that significantly increase misuse potential or exhibit low-level autonomous capabilities. ASL-4 and higher levels, not yet defined, will address more severe misuse potential and AI autonomy.
The ASL framework’s detailed criteria and safety measures are outlined in the main document. ASL-2 represents Anthropic’s current safety standards, aligning closely with their recent White House commitments. ASL-3 includes more stringent protocols, such as extremely strong security measures and a commitment to not deploying ASL-3 models if they pose misuse risks, even under extensive red-teaming by top experts. Future ASL-4 measures will address unsolved research challenges, requiring innovative assurance methods to prevent catastrophic AI behaviors.
Anthropic aims to balance targeting catastrophic risks with promoting beneficial AI applications and safety progress. The ASL system necessitates pausing the training of more advanced models if safety compliance cannot be maintained. This approach incentivizes solving safety issues to progress further, leveraging the most potent models from previous ASL levels to develop safety features for subsequent levels. This could create a competitive environment where solving safety problems becomes pivotal.
From a business perspective, the RSP will not impact the current use of Claude or the availability of Anthropic’s products. The policy is similar to pre-market safety testing in industries like automotive or aviation, aiming to ensure product safety before market release, ultimately benefiting customers.
Approved by Anthropic’s board, the RSP includes procedural safeguards to ensure evaluation integrity. Acknowledging the fast pace and uncertainties in AI development, Anthropic anticipates that rapid iteration and course correction will be essential for the RSP’s evolution.
Originally approved by Anthropic’s board, including consultations with the Long Term Benefit Trust, the policy is expected to be flexible to adapt to the fast-paced AI field. Anthropic designed the ASL system to benefit all stakeholders, helping policymakers, third-party nonprofit organizations, and other corporations make informed deployment decisions.
The development of the RSP has been supported by ARC Evals, whose expertise in AI risk assessment was crucial. Anthropic recognizes ARC Evals’ leadership in creating broader responsible scaling frameworks that inspired Anthropic’s approach.
Read the full Responsible Scaling Policy document here: https://www.anthropic.com/news/anthropics-responsible-scaling-policy