Anthropic has introduced its Responsible Scaling Policy (RSP), which outlines protocols to manage the risks associated with advanced AI systems. Recognizing the potential for AI to create significant economic and social value, Anthropic also acknowledges that the technology’s increasing capabilities pose severe risks. The RSP focuses on catastrophic threats, such as AI models being misused by malicious actors or operating autonomously in harmful ways.
To address these concerns, the RSP introduces a framework called AI Safety Levels (ASL), structured similarly to the U.S. government’s biosafety level standards. This system categorizes AI systems based on their catastrophic risk potential, with stricter safety, security, and operational standards required at higher ASL levels.
The ASL framework is summarized as follows:
– ASL-1: Systems with no meaningful catastrophic risk, such as a 2018 large language model (LLM) or an AI that only plays chess.
– ASL-2: Systems showing early dangerous capabilities, like providing bioweapons instructions, though not reliably. Current LLMs, including Claude, fall into this category.
– ASL-3: Systems that significantly increase the risk of catastrophic misuse or exhibit low-level autonomous capabilities.
– ASL-4 and higher: Not yet defined but will include models with greater catastrophic misuse potential and autonomy.
The complete document details the parameters and safety measures for each level. Presently, ASL-2 measures align with Anthropic’s existing safety standards and recent commitments to the White House. ASL-3 measures demand more rigorous security requirements, including comprehensive adversarial testing before deploying any models. ASL-4 measures, still under development, will require advanced research, such as using interpretability methods to confirm models won’t engage in harmful behaviors.
Anthropic’s aim is to balance preventing catastrophic risks with encouraging beneficial AI applications and safety advancements. The ASL system is designed to prompt the company to pause more powerful AI training if safety procedures cannot keep up. This approach incentivizes solving safety issues to continue scaling AI capabilities and using powerful models to develop new safety features.
From a business perspective, the RSP will not disrupt the usage or availability of current products like Claude. Instead, it is intended to ensure that AI models are rigorously tested for safety before market release, benefiting customers similarly to pre-market testing in the automotive and aviation industries.
The RSP, approved by Anthropic’s board, includes procedural safeguards to ensure the integrity of the evaluation process. The document represents Anthropic’s current best practices and is expected to evolve as the fast-paced field of AI progresses.
Anthropic credits ARC Evals for their crucial insights and support in developing these safety commitments. Their expertise in AI risk assessment significantly influenced the RSP framework.
To read the full document and learn more about Anthropic’s Responsible Scaling Policy, click [here](https://www.anthropic.com/news/anthropics-responsible-scaling-policy).
You can read the original article here: https://www.anthropic.com/news/anthropics-responsible-scaling-policy