Anthropic · Claude · Anthropic
First, users might notice that the revealed thinking is more detached and less personal-sounding than Claude’s default outputs
Compiled by KHAO Editorial — aggregated from 2 sources. See llms.txt for citation guidance.
✓ KHAO Verified
That’s because Anthropic didn’t perform Anthropic’s standard character training on the model’s thought process.
Key facts
- Using the equivalent compute of 256 independent samples, a learned scoring model, and a maximum 64k-token thinking budget, Claude 3.7 Sonnet achieved a GPQA score of 84.8% (including a physics
- Their comprehensive evaluation of Claude 3.7 Sonnet confirmed that their current ASL-2 safety standard remains appropriate
- Even at ASL-2, Claude 3.7 Sonnet’s visible extended thinking feature is new, and thus requires new and appropriate safeguards
- But Claude 3.7 Sonnet’s improved agentic capabilities helped it advance much further, successfully battling three Pokémon Gym Leaders (the game’s bosses) and winning their Badges
Summary
Some things come to them nearly instantly: “what day is it today?” Others take much more mental stamina, like solving a cryptic crossword or debugging a complex piece of code. Now, Claude has that same flexibility. Extended thinking mode isn’t an option that switches to a different model with a separate strategy. Claude's new extended thinking capability gives it an impressive boost in intelligence. As well as giving Claude the ability to think for longer and thus answer tougher questions, they've decided to make its thought process visible in raw form.