Claude · Anthropic
Fable 5’s classifiers are not intended to block these, and any blocks that do occur are likely to be false
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
The following are topics that have overlap with cybersecurity, but that are out of the scope of Anthropic’s cybersecurity classifiers.
Key facts
- For Fable 5, they expect to block these types of actions until they have better controls to limit access to known good actors
- For Claude Fable 5, they aim to block high- uplift vulnerability finding
- Here, they provide a detailed list of the types of harms Fable 5’s classifiers are, and are not, designed to prevent
- They've also launched a HackerOne program where security researchers can submit potential cyber jailbreaks they discover in Fable 5 for their review
Summary
First, they provide more information on the cybersecurity safeguards —specifically, the safety classifiers —that they launched with the model. Second, they lay out an early draft version of their proposed AI jailbreak severity framework, on which they've been working with their Glasswing partners. Jailbreaks vary in severity: sometimes they only unblock minor undesirable behaviors, and sometimes they unblock a wide range of harmful outputs, making a model much more dangerous. What they're sharing today reflects their current thinking. The team believe that by working together, they can establish a standard that enables the defensive uses of this technology while preventing its misuse.