Anthropic · Dario Amodei · MIT Technology Review
What Anthropic’s latest AI discovery does—and doesn’t—show
Compiled by KHAO Editorial — aggregated from 1 source + 2 references discovered via search. See llms.txt for citation guidance.
◎ Multiple-sources
This story originally appeared in The Algorithm, their weekly newsletter on AI.
Key facts
- Anthropic—currently the world’s most valuable AI company, with a nearly $1 trillion valuation—has a reputation for publishing strange and heady research
- Anthropic’s CEO, Dario Amodei, has said they won’t be able to control LLMs fully unless they learn more about how they work
- Describing AI models with terms borrowed from psychology and neuroscience can make their behavior seem more sophisticated than they might otherwise judge it
- Senior editor Will Douglas Heaven, aside from having a PhD in computer science, has spent a lot of time digging into what they can say about how AI models work
Summary
Anthropic—currently the world’s most valuable AI company, with a nearly $1 trillion valuation—has a reputation for publishing strange and heady research. One niche that Anthropic spends more time and money on than other AI companies is called mechanistic interpretability, which means looking inside the complex math of an AI model to learn why it comes up with one particular output and not another. That’s why, when Anthropic announced last week that it had found a new window into its models’ “internal thoughts” as they reason through answers, there was one colleague the reporter had to talk to. What did Anthropic learn here, exactly? Anthropic has been trying to understand how large language models (LLMs) work for a few years now.