Anthropic has introduced its Responsible Scaling Policy (RSP), outlining protocols to manage the risks associated with developing advanced AI systems. This policy aims to balance the economic and social benefits of AI with the potential for catastrophic hazards, such as misuse by malicious actors or unintended autonomous actions.
The core of the RSP is the AI Safety Levels (ASL) framework, which draws inspiration from the U.S. government’s biosafety level standards. This framework categorizes AI systems based on their risk levels, demanding stringent safety and security measures as the potential for catastrophic risk increases.
ASL-1 encompasses systems with negligible catastrophic risk, such as a 2018 language model or an AI designed solely for playing chess. ASL-2 covers systems that exhibit early signs of dangerous capabilities but are not yet reliably harmful. Current large language models, including Claude, fall into this category. ASL-3 applies to systems that significantly elevate misuse risks compared to non-AI baselines or exhibit low-level autonomous functions.
The policy remains open-ended for ASL-4 and above, as these will involve models with advanced capabilities that are not yet fully understood. The establishment of ASL-4 measures will require breakthroughs in safety and assurance techniques not currently available.
Anthropic’s ASL-2 standards align closely with recent commitments made to the White House, while ASL-3 standards will demand extensive research to implement. These may include robust security protocols and a pledge not to deploy ASL-3 models if they pose any risk of catastrophic misuse, validated through adversarial testing by top-tier red teams.
The ASL system is designed to encourage the development of beneficial applications while addressing potential risks. It necessitates halting the training of more powerful models if safety compliance can’t keep pace with AI scaling. However, it also incentivizes solving safety challenges to enable further advancements.
From a customer perspective, the RSP aims to ensure that products, such as Claude, continue to be safe and available. This process is likened to pre-market safety testing in industries like automotive or aviation, ultimately benefiting end-users.
The RSP has been formally approved by Anthropic’s board and will be subject to future board approvals after consulting with the Long Term Benefit Trust. The policy is an evolving framework, and adjustments will be made as AI technology and its landscape rapidly progress.
Anthropic credits ARC Evals for their vital input in developing the RSP, particularly their expertise in evaluating AI autonomy risks. ARC Evals’ broader Responsible Scaling Policy framework served as a key inspiration for this initiative.
You can read the original article here: https://www.anthropic.com/news/anthropics-responsible-scaling-policy