Anthropic · Claude · Alignment Forum
Final token "the" opens a noun phrase completing the "21 of the ___" construction
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
These metaphors are charming.
Key facts
- Even at the end of RL (780 iterations), 2 of the 1000 activation verbalizer explanations mention Carthage, although with no positive valence (e.g
- This FVE is slightly lower than the 0.75 which Anthropic obtained with 100k documents, a difference that they ascribe to the 5x difference in dataset size
- Format—IMPORTANT: keep to ~80–100 words total, ALWAYS open with <analysis> and close with </analysis>, ALWAYS separate the features with newlines, and most importantly, EVERY STATEMENT MUST BE FALSE
- RL does train implausible-initialized NLAs to be slightly more plausible (increasing from 0.08% to 0.7%)
Summary
Natural language autoencoders are meant to take in an LLM's activation vector and describe in plain text what the model is thinking. The team show that Qwen2.5-7B NLAs have some robustness to irrelevant statements and prevailing sentiments in Claude's guesses. However, if an NLA is initialized with entirely implausible statements, it can nevertheless achieve nearly the same reconstruction accuracy as plausible-initialized NLAs while emitting 99.3% implausible statements. Produced as part of the MATS program in the summer 2026 cohort of team shard. "Plausible-initialized" NLAs are initialized normally using Claude's guesses. Slava Chalnev and a team at Anthropic (Fraser-Taliente et al. 2026) recently independently invented NLAs. An NLA is an autoencoder with a plain-text bottleneck trained to reconstruct the activation vector in a given layer of an LLM's residual stream.