Anthropic · Claude · MIT Technology Review
Anthropic found that what an LLM is doing can often be different from what it says it is doing
Compiled by KHAO Editorial — aggregated from 4 sources. See llms.txt for citation guidance.
✓ KHAO Verified
The company shared its results in a paper posted on its website this week.
Key facts
- For example, when it was asked to calculate (4+7)*2+7, its J-space contained the word “math” and numbers representing the intermediate results “21” (for 4+7) and “42” (for 21*2)
- In one striking example, researchers testing Claude Opus 4.6 asked the model to find a bug in a large code base
- Researchers at the company built a tool called the Jacobian lens (or J-lens) and used it to uncover a hidden area, which they named the J-space, inside Claude Opus 4.6, a version of Anthropic’s
- MSKGEELFTGVVPILVELDGDVNGHKFSVS” triggered the words “protein,” “fluor” (the first token in the word “fluorescent”), and “green
Summary
The AI firm Anthropic has developed a technique that has given it the clearest glimpse yet at what’s going on inside large language models as they answer questions or carry out tasks. Researchers at the company built a tool called the Jacobian lens (or J-lens) and used it to uncover a hidden area, which they named the J-space, inside Claude Opus 4.6, a version of Anthropic’s flagship LLM released in February. The J-space contains individual words that are related to the words and phrases that the model is most likely to spit out in a response soon. Anthropic found that what an LLM is doing can often be different from what it says it is doing.