← Back to KHAO

The Information · Google ·

Investigating this risk, Engels et al. surprisingly surfaced that DiffusionGemma scores similar monitorability to Gemma

2 min read

Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.

◌ Single Source

GPQA truncation failure modes.

Thus somehow the information in seemed to have been essential, conflicting with the results of high monitorability.

Key facts

Summary

Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. Still, they also find some rare case studies where the distribution vector is load-bearing computationally, i.e. where top-1 projection would be detrimental. Apart from model behavior, they also examined how interpretability techniques carry over to DiffusionGemma, including probes, steering, and J-lens. Overall, this supports the paper's conclusion that DiffusionGemma remains highly monitorable, while nevertheless showing that there are cases where models can learn to use vector-valued information. Large language models arrive at answers to complex questions through chains-of-thought. DiffusionGemma is a particular model whose architecture allows latent reasoning, by passing a vector encoding a probability distribution between steps.

Read full article at Alignment Forum →

#The Information #Google