Anthropic has published its Responsible Scaling Policy (RSP), a comprehensive set of protocols aimed at managing the risks associated with increasingly capable AI systems. These guidelines seek to balance the creation of economic and social value with the potential risks AI models pose, particularly catastrophic risks like deliberate misuse for bioweapon creation or autonomous actions leading to large-scale harm.
The RSP introduces a framework called AI Safety Levels (ASL), inspired by the US government’s biosafety level standards for handling hazardous biological materials. The ASL framework sets progressively stringent safety, security, and operational standards based on the potential catastrophic risk of an AI model.
At the most basic level, ASL-1 applies to systems posing no substantial catastrophic risk, such as a 2018 AI model or one that solely plays chess. ASL-2 covers systems with initial signs of dangerous capabilities, such as providing unreliable instructions on bioweapon creation. Current Large Language Models (LLMs), including Anthropic’s Claude, fall into this category. ASL-3 is for systems that significantly elevate the risk of misuse or exhibit low-level autonomous capabilities. ASL-4 and higher are not yet defined but will involve more intricate safety measures as the technology evolves.
Anthropic’s RSP is designed to encourage both effective targeting of catastrophic risks and the incentivization of beneficial AI applications and safety advancements. The ASL framework mandates that more powerful models are only trained if the necessary safety protocols are in place. This approach aims to resolve safety issues proactively, using current advanced models to develop safety features for future, more powerful models.
From a business standpoint, the RSP will not impact the current uses of the Claude AI or disrupt product availability. Similar to pre-market testing in automotive or aviation industries, the RSP seeks to assure product safety before market release, ultimately benefiting customers. The policy has been ratified by Anthropic’s board and future changes will require board approval, ensuring procedural safeguards are in place.
Anthropic acknowledges that the RSP is a work in progress, reflecting the fast-paced and uncertain nature of the AI field. Adjustments and updates to the policy are expected as AI technology continues to evolve.
The development of the RSP involved significant contributions from ARC Evals, whose expertise in AI risk assessment was instrumental. ARC Evals’ broader framework also inspired Anthropic’s approach.
Anthropic’s RSP offers a pioneering approach to AI safety and risk management, aiming to set a standard that could influence broader industry practices. The full details of the policy are available for review and provide valuable insights for policymakers, nonprofit organizations, and other companies facing similar AI deployment challenges.
You can read the original article here: https://www.anthropic.com/news/anthropics-responsible-scaling-policy