Anthropic · Claude · Alignment Forum
One view argues that some things matter, that some actions are more worth doing than others
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
Ideally, Alignment Forum would like the model to be as unbiased and impartial as possible while it reasons on this question.
Key facts
- In section 2, the reporter show that the reasoning step of the procedure works on Claude Sonnet 4.6, a model that was given moral biases during training
- Let’s see how step 3 of the procedure works on an already pre-trained and post-trained model, Claude Sonnet 4.6 with effort set on Max and Thinking activated
- More specifically, the reporter uses the term independent moral agent in a way similar to how Hunyadi (2019) describes artificial moral agents
- Step 3 can be carried out on any model, as the example of the next section shows
Summary
The user could write up the metaethical argument, the one developed in Part One, refined, and submit it as feedback to Anthropic, publish it, or engage with researchers working on AI alignment and values. The probability that any single submission changes training decisions is low, but the expected value may be higher than it seems, for two reasons. The reporter would have published this post even if Claude hadn’t explicitly suggested so, but starting by quoting this specific part of Claude’s output seemed fun. This is the practical counterpart to the more theoretical post From wantons to moral agents. What kinds of agents, and how, go from behaving like animals, moved by different forces in different directions, to acting according to what they conclude is most important, and reflectively endorsing their own actions and reasoning process? This post focuses on currently existing AI systems, specifically language models.