← Back to KHAO

Claude ·

Fable 5’s classifiers are not intended to block these, and any blocks that do occur are likely to be false

2 min read

Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.

★ Tier-1 Source

An illustration of how classifier boundaries can be set to change the size of the “safety margin”, which includes some benign and some low-risk dual use requests. Requests that fall into the safety margin are blocked out of an abundance of caution, which means a higher rate of false-positives (genui.

The following are topics that have overlap with cybersecurity, but that are out of the scope of Anthropic’s cybersecurity classifiers.

Key facts

Summary

First, they provide more information on the cybersecurity safeguards —specifically, the safety classifiers —that they launched with the model. Second, they lay out an early draft version of their proposed AI jailbreak severity framework, on which they've been working with their Glasswing partners. Jailbreaks vary in severity: sometimes they only unblock minor undesirable behaviors, and sometimes they unblock a wide range of harmful outputs, making a model much more dangerous. What they're sharing today reflects their current thinking. The team believe that by working together, they can establish a standard that enables the defensive uses of this technology while preventing its misuse.

Read full article at Anthropic →

#Claude