Anthropic Unveils AI Responsible Scaling Policy to Mitigate Catastrophic Risks.

Anthropic has unveiled its Responsible Scaling Policy (RSP) to manage the risks associated with increasingly capable AI systems. The policy includes a series of technical and organizational protocols aimed at addressing catastrophic risks—those where an AI model could cause large-scale devastation. These could arise from either the deliberate misuse of models, such as by terrorists creating bioweapons, or from autonomous actions that diverge from the designer’s intent.

The RSP introduces AI Safety Levels (ASL), a framework inspired by the U.S. government’s biosafety level standards for dangerous biological materials. This framework categorizes AI models by their potential for catastrophic risk and mandates specific safety, security, and operational standards for each level.

ASL-1 encompasses systems with no meaningful catastrophic risk, like a 2018 language model or an AI that only plays chess. ASL-2 includes systems showing initial signs of dangerous capabilities, such as providing crude instructions for bioweapons without reliability. Most current large language models, including Anthropic’s Claude, fall into this category. ASL-3 systems significantly increase the risk of misuse or exhibit low-level autonomous capabilities. ASL-4 and higher levels are not yet defined but will involve more stringent safety measures as AI capabilities advance.

The ASL system requires rigorous safety demonstrations, especially as models become more capable. For example, while ASL-2 measures align closely with recent White House commitments on AI safety, ASL-3 will necessitate stronger standards, including advanced security and a commitment not to deploy models that show significant misuse potential under rigorous testing.

The policy aims to balance mitigating catastrophic risks while promoting beneficial AI applications and safety research. This dynamic approach might temporarily pause the training of powerful models if safety procedures can’t keep up. However, it incentivizes solving safety issues to unlock further advancements, potentially fostering a “race to the top” in safety innovation among AI developers.

Business operations will not be affected by the RSP. Instead, it parallels safety protocols in industries like automotive or aviation, where rigorous testing precedes market release. This approach benefits customers by ensuring product safety.

Anthropic’s board has formally approved the RSP, and any changes will require further board approval, involving consultations with the Long Term Benefit Trust. The company acknowledges that rapid iteration and course correction will be necessary given the fast-paced evolution of AI.

The full document of the Responsible Scaling Policy is available online and is intended to inspire policymakers, nonprofit organizations, and other companies making similar deployment decisions. Anthropic thanks ARC Evals for their expertise in shaping the RSP commitments, particularly in the area of AI risk assessment.

You can read the original article here: [Anthropics Responsible Scaling Policy](https://www.anthropic.com/news/anthropics-responsible-scaling-policy)

Leave a Reply

Discover more from Innovation Era

Subscribe now to keep reading and get access to the full archive.

Continue reading