DeepSeek · Claude · GPT · Gemini · Decrypt
DeepSeek's New Model Nearly Matches GPT-6 Astra on Design—at 1.4% of the Cost
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
OpenDesign, the company behind the benchmark site OpenDesign Arena, ran 13 AI models through the same batch of design tasks this week.
Key facts
- Claude Fable 5.1 came in at 80.3, took 12.8 minutes, and cost $3.66
- On that scale, GPT-6 Astra averaged 82.7 points, taking 11.1 minutes and $1.61 per finished design
- But DeepSeek's newest model, V4.1 Flash, reached 98% of that top score while charging about 1.4% of the top price
- Every other model OpenDesign tested—Grok 4.6, Qwen 3.8-Max, Kimi K3, GLM-5.3 Flash, and Gemini 3.8 Flash among them—scored lower than DeepSeek V4.1 Flash and cost more to run
Summary
OpenDesign Arena scored DeepSeek V4.1 Flash at 81.2 out of 100 on real-world design tasks, 98% of GPT-6 Astra's 82.7, while charging $0.023 per finished design against Astra's $1.61. Of the 13 models tested—including Claude Fable 5.1, Grok 4.6, and Qwen 3.8-Max—11 scored lower than DeepSeek's model and cost more to run. DeepSeek's technical paper for V4.1 Flash shows the model activates 8 billion of its 552 billion parameters to read a prompt, the design choice behind its low price. OpenDesign Arena scores models on everyday design work—building web apps, dashboards, mobile screens, and landing pages—out of 100 points.