← Back to KHAO

Blue Origin ·

Spotting Behaviors to investigate

2 min read

Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.

◌ Single Source

updated figure.

This work was largely done during Neel Nanda's MATS 10.0 Exploration Phase.

Key facts

Summary

J Rosser and Dohun Lee are co-first authors for this post with equal contribution. A natural assumption is that they can control what a LLM learns during training by controlling the data. Surprisingly, this didn't work! The team handpicked a set of behaviors where the SFT model differed from the mid-train. The team first create a lower cost “speed-run” version of the full SFT set up.

Read full article at Alignment Forum →

#Blue Origin