AI Agent · stepfun.com
Step 5 Preview: Advancing the Pareto Frontier
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
The evaluations above highlight a subset of Step 5 Preview’s capabilities.
Key facts
- The team gave Step 5 Preview 24 hours to optimize an MLA GPU kernel from scratch on an NVIDIA H100 GPU, with a head dimension of 512 and a production shape of batch size 1, 64 heads, and 8,192 tokens
- DeepSWE v1.1 evaluations use the SWE-agent harness with temperature=1.0 and top_p=0.95
- Without any Pokémon-specific optimization, Step 5 Preview has so far sustained progress for more than 3,000 turns and 6 million tokens of interaction
- The team also evaluated Step 5 Preview on FrontierFinance, an external benchmark covering six investment use cases through 220 expert-crafted questions and 11,543 evaluation criteria
Summary
The Pareto frontier marks the best trade-offs between intelligence and cost. Across public and internal evaluations, Step 5 Preview performs strongly across software engineering, agentic tasks, professional knowledge work, and finance. GDPval-AA v2 scores are based on the latest results from Artificial Analysis, as of Sep. 19, 2026. DeepSWE v1.1 evaluations use the SWE-agent harness with temperature=1.0 and top_p=0.95. Step 5 Preview's coding capabilities cover a broad range of development work, from software engineering and visual applications to programmable hardware.