Anthropic has introduced its Responsible Scaling Policy (RSP), which sets out technical and organizational protocols designed to manage the risks associated with increasingly capable AI systems. With AI models becoming more powerful and influential, there is potential for both major economic and social benefits, as well as significant risks. The RSP specifically addresses catastrophic risks, which are those resulting in vast devastation from AI misuse or unintended autonomous actions.
The policy includes the AI Safety Levels (ASL), a framework inspired by the U.S. government’s biosafety level standards for handling hazardous biological materials. The ASL system categorizes AI risks into different levels, requiring more rigorous safety and security measures as the risk level increases. For instance, ASL-1 encompasses systems with no catastrophic risk, while ASL-2 includes models showing early dangerous capabilities but aren’t yet practical for harmful uses. Current AI models, including Claude, are classified under ASL-2.
ASL-3 signals a higher risk of catastrophic misuse or low-level autonomous capabilities. ASL-4 and higher levels are not yet defined but are expected to deal with substantial autonomy and misuse potential. The criteria and safety measures for each ASL level are elaborated in the full policy document. Presently, ASL-2 measures align with Anthropic’s current standards and commitments, such as those made to the White House. ASL-3 will require more stringent security measures and a commitment to prevent model deployment if they show catastrophic misuse risks during adversarial testing.
Anthropic aims to balance mitigating catastrophic risks with promoting beneficial AI applications and advancing safety research. The ASL system motivates the company to resolve safety issues to continue scaling AI capabilities. This framework, if adopted widely, could encourage a competitive drive to enhance AI safety.
From a business standpoint, the RSP will not change current uses of Claude or affect product availability. It is likened to pre-market testing in industries like automotive or aviation, ensuring product safety before public release. The policy has been approved by Anthropic’s board and any changes will undergo board evaluation following consultative procedures.
Anthropic acknowledges that the RSP represents an early iteration and anticipates the need for rapid adaptation due to the fast-paced and unpredictable nature of the AI field. Further details and procedural safeguards are available in the full RSP document, intended to inspire policymakers and other companies facing similar AI deployment challenges.
The company expressed gratitude to ARC Evals for their expertise, which was crucial in developing the RSP commitments and evaluation procedures. The ARC Responsible Scaling Policy framework significantly influenced Anthropic’s approach.
You can read the original article here: https://www.anthropic.com/news/anthropics-responsible-scaling-policy