Finally, the charts below show how Claude Opus 4.6 performs on a variety of benchmarks that assess its software engineering
·2 min read
Compiled by KHAO Editorial
— aggregated from 2 sources + 4 references discovered via search.
See llms.txt for citation guidance.
★ Tier-1 Source
These intelligence gains do not come at the cost of safety.
Key facts
On GDPval-AA —an evaluation of performance on economically valuable knowledge work tasks in finance, legal, and other domains 2 —Opus 4.6 outperforms the industry’s next-best model (OpenAI’s GPT-5.2)
Pricing remains the same at $5/$25 per million tokens; for full details, see their pricing page
[3] This translates into Claude Opus 4.6 obtaining a higher score than GPT-5.2 on this eval approximately 70% of the time (where 50% of the time would have implied parity in the scores)
And, in a first for their Opus-class models, Opus 4.6 features a 1M token context window in beta 1
Summary
The new Claude Opus 4.6 improves on its predecessor’s coding skills. Opus 4.6 can also apply its improved abilities to a range of everyday work tasks: running financial analyses, doing research, and using and creating documents, spreadsheets, and presentations. The model’s performance is state-of-the-art on several evaluations. As they show in their extensive system card, Opus 4.6 also shows an overall safety profile as good as, or better than, any other frontier model in the industry, with low rates of misaligned behavior across safety evaluations. In Claude Code, you can now assemble agent teams to work on tasks together.