Anthropic’s New AI Safety Levels: Ensuring Responsible AI Progress.

Anthropic has unveiled its Responsible Scaling Policy (RSP), a new framework designed to manage the risks associated with increasingly capable AI systems. The company aims to address the catastrophic risks posed by advanced AI, such as misuse by malicious actors or systems acting autonomously against their designers’ intent.

The RSP introduces a structured approach known as AI Safety Levels (ASL), inspired by biosafety levels used in handling hazardous biological materials. This system will set safety, security, and operational standards based on a model’s potential for catastrophic risk. The higher the ASL level, the stricter the safety guidelines.

At the base, ASL-1 includes systems with no significant catastrophic risk, such as a 2018 language model or a chess-playing AI. ASL-2 includes models showing early signs of dangerous capabilities, like instructions for bioweapon creation, without being fully reliable. Current language models, including Claude, fall into this category. ASL-3 escalates to systems presenting a greater risk of catastrophic misuse or autonomous behavior. ASL-4 and higher levels are not yet defined but will handle more substantial risks and autonomous capabilities.

The RSP document details the criteria and safety measures for each ASL level. ASL-2 standards overlap with Anthropic’s recent commitments to the White House, while ASL-3 will require rigorous security measures and red-teaming efforts. ASL-4 measures, pending development, will address today’s unsolved safety challenges.

Anthropic designed the ASL system to balance effectively managing catastrophic risks with incentivizing beneficial applications and safety advancements. The system encourages solving safety issues to unlock further AI capabilities rather than merely pausing developments.

From a business standpoint, the RSP will not disrupt current uses of Claude or availability of Anthropic’s products. Instead, it is akin to pre-market testing in the automotive or aviation industries, aiming to ensure rigorous product safety before market release.

The RSP has been formally approved by Anthropic’s board, with changes requiring board approval and consultations with the Long Term Benefit Trust. The company acknowledges that rapid iteration and course correction will be necessary given AI’s fast pace and uncertainties.

The full Responsible Scaling Policy document is available for policymakers, nonprofit organizations, and other companies facing similar AI deployment challenges. Anthropic credits ARC Evals for their essential insights and expertise in developing the RSP commitments.

You can read the original article here: https://www.anthropic.com/news/anthropics-responsible-scaling-policy

Leave a Reply

Discover more from Innovation Era

Subscribe now to keep reading and get access to the full archive.

Continue reading